Practical Guide to Homology Modeling: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
Eric Martz (talk | contribs)
Line 56: Line 56:
#Push the [[Image:Rcsb-search-button.png]] button to run the search.
#Push the [[Image:Rcsb-search-button.png]] button to run the search.
#Scroll down to see the list of hits.
#Scroll down to see the list of hits.
#At the top of the list, change <font color='red'>Display Results as</font> to <font color='red'>'''Polymer Entities'''</font>. Then push [[Image:Rcsb-search-button.png]] again. <font color='red'>''This is crucial''</font> because it displays the identity percentages and alignments for the hits.
#At the top of the list, change <font color='red'>Display Results as</font> to <font color='red'>'''Polymer Entities'''</font>. Then push [[Image:Rcsb-search-button.png]] again. <font color='red'>''This is crucial''</font> because it displays the identity percentages and alignments for the hits. It should be the default!
#The best hits will be listed first, starting below “Showing 1-25 of NNN”.  Notice that each hit starts with a large, bold PDB ID.
#The best hits will be listed first.  Notice that each hit starts with a large, bold PDB ID.


For each hit, notice the “Identities” above the sequence alignment box. The denominator tells you the length of the sequence alignment. The percentage tells you the sequence identity of the alignment.
For each hit, notice the '''Sequence Identity %''' above the sequence alignment box.


For example, “Identities: 355/1045 (34%)” means that 1,045 residues of your query sequence align to the hit with 34% sequence identity (355 identical residues in the alignment). Knowing that my query had length 1,170 residues, I can see that this potential template for a homology model would enable me to model 1,045/1,170 = 89% of my query sequence. Quite often the alignment would span a much smaller portion of the full-length sequence.
Also notice the '''Region''' range, which tell you how many of your query residues align with the hit. Compare this to the full length of your query sequence. 
 
<!--For example, “Identities: 355/1045 (34%)” means that 1,045 residues of your query sequence align to the hit with 34% sequence identity (355 identical residues in the alignment). Knowing that my query had length 1,170 residues, I can see that this potential template for a homology model would enable me to model 1,045/1,170 = 89% of my query sequence. Quite often the alignment would span a much smaller portion of the full-length sequence.


<font color='red'>'''BEWARE!''' If you forgot to set <i>Mask Low Complexity</i> to NO:</font> The sequence identity percentage may be '''underestimated''' at rcsb.org. This happens when rcsb.org deems segments of the query sequence to be of low complexity. Such segments are marked with X’s in the sequence alignment, and excluded from the calculation of sequence identity. For example, for Saccharomyces gal4 (UniProt P04386), for the top hit (3coq), rcsb.org reports “Identities: 71/89 (80%)”, while in fact the sequence identity is 100%. Note this in the sequence alignment at rcsb.org:
<font color='red'>'''BEWARE!''' If you forgot to set <i>Mask Low Complexity</i> to NO:</font> The sequence identity percentage may be '''underestimated''' at rcsb.org. This happens when rcsb.org deems segments of the query sequence to be of low complexity. Such segments are marked with X’s in the sequence alignment, and excluded from the calculation of sequence identity. For example, for Saccharomyces gal4 (UniProt P04386), for the top hit (3coq), rcsb.org reports “Identities: 71/89 (80%)”, while in fact the sequence identity is 100%. Note this in the sequence alignment at rcsb.org:
[[Image:Seq-algn-lo-complexity.png|center]]
[[Image:Seq-algn-lo-complexity.png|center]]


The 18 residues marked X were not included in the identity calculation. In contrast, when the same sequence search is performed at [http://www.ebi.ac.uk/pdbe PDB-Europe], 100% sequence identity is reported. However, other aspects of the report at PDB-Europe are less satisfactory (e.g. the length of the alignment is not stated; the sequences are not numbered) and hence we recommend using rcsb.org despite its misleading sequence identity percentages.
The 18 residues marked X were not included in the identity calculation. In contrast, when the same sequence search is performed at [http://www.ebi.ac.uk/pdbe PDB-Europe], 100% sequence identity is reported. However, other aspects of the report at PDB-Europe are less satisfactory (e.g. the length of the alignment is not stated; the sequences are not numbered) and hence we recommend using rcsb.org despite its misleading sequence identity percentages.-->


== Are parts (or all) of the query protein intrinsically disordered? ==
== Are parts (or all) of the query protein intrinsically disordered? ==