Structure superposition tools: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
m Structural superposition tools moved to Structure superposition tools: Further improvement in title.
Eric Martz (talk | contribs)
No edit summary
Line 1: Line 1:
''Structural superposition'' refers to the optimal superposition, yielding the closest fit, in three dimensions, between two or more molecular models. It is sometimes called ''structural alignment'', but that term more often denotes a sequence alignment guided by a structural superposition. In the case of proteins, structural superposition is often performed without reference to the sequences of the proteins. When the models superpose closely, it suggests evolutionary and functional relationships that may not be discernable from sequence comparisions<ref>PMID: 10686110</ref>.  
''Structure superposition'' refers to the optimal superposition, yielding the closest fit in three dimensions, between two or more molecular models. It is sometimes called ''structure alignment'', but that term is easily confused with a sequence alignment guided by a structure superposition. In the case of proteins, structure superposition is often performed without reference to the sequences of the proteins. When the models superpose closely, it suggests evolutionary and functional relationships that may not be discernable from sequence comparisions<ref>PMID: 10686110</ref>.  


The purpose of this article is to help in choosing a server or software package for performing structural superposition. Characteristics of structural superposition servers and software packages are listed, along with results of testing with a few examples.
The purpose of this article is to help in choosing a server or software package for performing structure superposition. Characteristics of structure superposition servers and software packages are listed, along with results of testing with a few examples.


There are two common applications of structural superposition servers:
There are two common applications of structure superposition servers:
# '''Pairwise superposition'''. All servers listed below enable you to upload two 3D models (or specify them from the [[PDB]]) and generate a structural superposition.
# '''Pairwise superposition'''. All servers listed below enable you to upload two 3D models (or specify them from the [[PDB]]) and generate a structure superposition.
# '''Structure neighbors'''. Some servers (notably [[#Dali|Dali]],  [[#FATCAT|FATCAT]], [[#VAST|VAST]] and [[#TopSearch|TopSearch]]) enable you to upload one 3D model (or specify one in the [[PDB]]) and generate a list of the closest structures in the [[PDB]], based on pairwise structural superpositions between your query structure versus each structure in the [[PDB]].
# '''Structure neighbors'''. Some servers (notably [[#Dali|Dali]],  [[#FATCAT|FATCAT]], [[#VAST|VAST]] and [[#TopSearch|TopSearch]]) enable you to upload one 3D model (or specify one in the [[PDB]]) and generate a list of the closest structures in the [[PDB]], based on pairwise structure superpositions between your query structure versus each structure in the [[PDB]].


Wikipedia offers a [http://en.wikipedia.org/wiki/Structural_alignment_software list of structural superposition software packages] and an overview of [http://en.wikipedia.org/wiki/Structural_alignment structural superposition]. Hasegawa and Holm reviewed structural superposition methods in 2009<ref>PMID: 19481444</ref>.
Wikipedia offers a [http://en.wikipedia.org/wiki/Structural_alignment_software list of structure superposition software packages] and an overview of [http://en.wikipedia.org/wiki/Structural_alignment structure superposition]. Hasegawa and Holm reviewed structure superposition methods in 2009<ref>PMID: 19481444</ref>.


==Evaluating Structural Superpositions==
==Evaluating Structure Superpositions==
The structural differences between two optimally superposed models are usually measured as the [http://en.wikipedia.org/wiki/RMSD Root Mean Square Deviation] ('''RMSD''') between the superposed alpha-carbon positions (excluding deviations from the non-superposed positions). To provide a frame of reference for RMSD values, note that up to 0.5 Å RMSD of alpha carbons occurs in independent determinations of the same protein<ref name="chothia86">PMID: 3709526</ref>. Crystallographic models of proteins with about 50% sequence identity differ by about 1 &Aring; RMSD<ref name="chothia86" /><ref name="3dcrunch">PMID: 10865955</ref>. Deviations can be much larger for models determined by [[NMR]]<ref name="3dcrunch" />.
The structural differences between two optimally superposed models are usually measured as the [http://en.wikipedia.org/wiki/RMSD Root Mean Square Deviation] ('''RMSD''') between the superposed alpha-carbon positions (excluding deviations from the non-superposed positions). To provide a frame of reference for RMSD values, note that up to 0.5 Å RMSD of alpha carbons occurs in independent determinations of the same protein<ref name="chothia86">PMID: 3709526</ref>. Crystallographic models of proteins with about 50% sequence identity differ by about 1 &Aring; RMSD<ref name="chothia86" /><ref name="3dcrunch">PMID: 10865955</ref>. Deviations can be much larger for models determined by [[NMR]]<ref name="3dcrunch" />.


The statistical significance of a structural superposition, relative to a superposition of random sequence-nonredundant structures in the [[PDB]], is usually measured with a '''[http://en.wikipedia.org/wiki/Standard_score z-score]'''. The z-score is the distance, in standard deviations, between the observed superposition RMSD and the mean RMSD for random pairs of the same length, with the same or fewer gaps. Z-scores less than 2 are considered to lack statistical significance.
The statistical significance of a structure superposition, relative to a superposition of random sequence-nonredundant structures in the [[PDB]], is usually measured with a '''[http://en.wikipedia.org/wiki/Standard_score z-score]'''. The z-score is the distance, in standard deviations, between the observed superposition RMSD and the mean RMSD for random pairs of the same length, with the same or fewer gaps. Z-scores less than 2 are considered to lack statistical significance.


When the models being compared have substantial differences, and especially if they have multiple domains, more tolerant estimates of the closenss of fit have been employed, notably in [[CASP]]. One of these is the ''global distance test total score'', or [[Calculating_GDT_TS|GDT_TS]]. See also [[Theoretical models]].
When the models being compared have substantial differences, and especially if they have multiple domains, more tolerant estimates of the closenss of fit have been employed, notably in [[CASP]]. One of these is the ''global distance test total score'', or [[Calculating_GDT_TS|GDT_TS]]. See also [[Theoretical models]].


==Visualizing Structural Superpositions==
==Visualizing Structure Superpositions==
<applet size='400' frame='true' align='right' caption='Structural alignment of [[1fsz]] with [[1tub]].'
<applet size='400' frame='true' align='right' caption='Structural alignment of [[1fsz]] with [[1tub]].'
scene='Structural_alignment_tools/Dali_chains_ab_water/1' />
scene='Structural_alignment_tools/Dali_chains_ab_water/1' />
Structural superpositions are usually visualized as the superposed backbone traces of the models. The example at right shows the bacterial cell division protein <font color="#d80000"><b>FtsZ</b></font> ([[1fsz]]:A) superposed by [[#Dali|Dali]] with <!--e0b000--><font color="#d0a000"><b>mammalian tubulin</b></font> ([[1tub]]:A). Sequence identity in the structurally superposed regions is about 13%.
Structure superpositions are usually visualized as the superposed backbone traces of the models. The example at right shows the bacterial cell division protein <font color="#d80000"><b>FtsZ</b></font> ([[1fsz]]:A) superposed by [[#Dali|Dali]] with <!--e0b000--><font color="#d0a000"><b>mammalian tubulin</b></font> ([[1tub]]:A). Sequence identity in the superposed regions is about 13%.
*The non-superposed segments are white in the query (<font color="#d80000"><b>FtsZ</b></font>) and thin in the target (<font color="#d0a000"><b>tubulin</b></font>). This scene is available in [[#Dali|Dali]] except that the target color has been changed to make it more distinct from the red query. (<scene name='Structural_alignment_tools/Dali_chains_ab_water/1'>Restore initial scene</scene>.)
*The non-superposed segments are white in the query (<font color="#d80000"><b>FtsZ</b></font>) and thin in the target (<font color="#d0a000"><b>tubulin</b></font>). This scene is available in [[#Dali|Dali]] except that the target color has been changed to make it more distinct from the red query. (<scene name='Structural_alignment_tools/Dali_chains_ab_water/1'>Restore initial scene</scene>.)
*Because the superposition is about 300 residues long (and the protein chains are longer), it is hard to see details of this superposition in the complexity. Buttons below show 50-residue segments of the query (<font color="#d80000"><b>FtsZ</b></font>) and backbone for target  (<font color="#d0a000"><b>tubulin</b></font>) where the target &alpha; carbons are within 3.5 &Aring;. (The RMSD for this [[#Dali|Dali]] superposition is 3.2 &Aring;.)
*Because the superposition is about 300 residues long (and the protein chains are longer), it is hard to see details of this superposition in the complexity. Buttons below show 50-residue segments of the query (<font color="#d80000"><b>FtsZ</b></font>) and backbone for target  (<font color="#d0a000"><b>tubulin</b></font>) where the target &alpha; carbons are within 3.5 &Aring;. (The RMSD for this [[#Dali|Dali]] superposition is 3.2 &Aring;.)
Line 53: Line 53:
</jmolButton>
</jmolButton>
</jmol>
</jmol>
* This <scene name='Structural_alignment_tools/Morph_1fsz_1tub_a_fatcat/1'>morph of the superposition</scene> was generated by [[#FATCAT|FATCAT]], which reported 3.02 &Aring; RMSD for 298 structurally superposed residues, and 10.2% sequence identity for the structurally superposed residues. The morph shows the 334-residue sequence of the query (FtsZ) changing from the query conformation to the conformation of the superposed target (tubulin). It does not show the non-superposed loops of tubulin that can be seen as thin backbone traces in the initial scene above. The morph makes it easy to see that the core fold is stable, while the larger changes occur in surface loops.
* This <scene name='Structural_alignment_tools/Morph_1fsz_1tub_a_fatcat/1'>morph of the superposition</scene> was generated by [[#FATCAT|FATCAT]], which reported 3.02 &Aring; RMSD for 298 superposed residues, and 10.2% sequence identity for the superposed residues. The morph shows the 334-residue sequence of the query (FtsZ) changing from the query conformation to the conformation of the superposed target (tubulin). It does not show the non-superposed loops of tubulin that can be seen as thin backbone traces in the initial scene above. The morph makes it easy to see that the core fold is stable, while the larger changes occur in surface loops.


It is very helpful to color the target alpha carbons by deviation ("RMSD") from the query model: red indicates large deviations (poor superposition) while blue indicates small deviations (close superposition), with white indicating average superposition. The stand-alone programs [[#DeepView = Swiss-PDBViewer|DeepView = Swiss-PDBViewer]] and [[#PyMOL|PyMOL]] color superpositions by RMSD but the results cannot be easily exported to Jmol. Surprisingly, none of the servers listed below color their superpositions by deviation, except [[#Dali|Dali]]. Unfortunately, there is NO built-in way to color the superposition by RMSD in Jmol.
It is very helpful to color the target alpha carbons by deviation ("RMSD") from the query model: red indicates large deviations (poor superposition) while blue indicates small deviations (close superposition), with white indicating average superposition. The stand-alone programs [[#DeepView = Swiss-PDBViewer|DeepView = Swiss-PDBViewer]] and [[#PyMOL|PyMOL]] color superpositions by RMSD but the results cannot be easily exported to Jmol. Surprisingly, none of the servers listed below color their superpositions by deviation, except [[#Dali|Dali]]. Unfortunately, there is NO built-in way to color the superposition by RMSD in Jmol.
Line 59: Line 59:
==Conclusions==
==Conclusions==
===Protein structural superposition===
===Protein structural superposition===
There are several well-documented, easy to use servers and software packages that do an excellent job of sequence-independent structural superposition, described below.  They include
There are several well-documented, easy to use servers and software packages that do an excellent job of sequence-independent structure superposition, described below.  They include
* [[#CE|CE]] rigid superposition only (see Note* below).
* [[#CE|CE]] rigid superposition only (see Note* below).
* [[#Dali|Dali]] rigid superposition only. Jmol. Colors by ''structure conservation'' distinguishing closely superposed from poorly superposed segments.
* [[#Dali|Dali]] rigid superposition only. Jmol. Colors by ''structure conservation'' distinguishing closely superposed from poorly superposed segments.