Conservation, Evolutionary: Difference between revisions
From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs) →The ConSurf-DB Mechanism: describing ConSurf-DB Mechanism |
Eric Martz (talk | contribs) →The ConSurf-DB Mechanism: describing ConSurf-DB Mechanism |
||
| Line 50: | Line 50: | ||
The ConSurf DataBase server, [http://consurfdb.tau.ac.il ConSurf-DB]<ref name="consurfdb">PMID: 18971256</ref>, pre-calculates conservation levels for each amino acid in every protein chain in the [[Protein Data Bank]]. It went into service in 2008. Each chain is processed as follows. | The ConSurf DataBase server, [http://consurfdb.tau.ac.il ConSurf-DB]<ref name="consurfdb">PMID: 18971256</ref>, pre-calculates conservation levels for each amino acid in every protein chain in the [[Protein Data Bank]]. It went into service in 2008. Each chain is processed as follows. | ||
#The amino acid sequence of each protein chain is submitted to PSI-BLAST<ref>PSI-BLAST (Position Specific Iteration-BLAST) is an extension of the Basic Local Alignment Search Tool (BLAST) that is more sensitive at finding distantly related sequences. See [http://en.wikipedia.org/wiki/PSI-BLAST PSI-BLAST at Wikipedia] and [http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/psi1.html PSI-BLAST at NCBI].</ref> for collection of related sequences. | #A list of unique protein chains is extracted from the [[Protein Data Bank]]. Chains shorter than 30 amino acids are not processed because they do not contain enough information for reliable phylogenetic tree construction. Non-standard residues are converted to the closest standard amino acids. Chains with more than 15% non-standard residues are not processed. | ||
#The amino acid sequence of each protein chain is submitted to PSI-BLAST<ref>PSI-BLAST (Position Specific Iteration-BLAST) is an extension of the Basic Local Alignment Search Tool (BLAST) that is more sensitive at finding distantly related sequences. See [http://en.wikipedia.org/wiki/PSI-BLAST PSI-BLAST at Wikipedia] and [http://www.ncbi.nlm.nih.gov/Education/BLASTinfo/psi1.html PSI-BLAST at NCBI].</ref> for collection of related sequences from UniprotKB/Swiss-Prot<ref>From [http://www.uniprot.org/help/uniprotkb UniProtKB help]: "UniProtKB/Swiss-Prot (reviewed) is a high quality manually annotated and non-redundant protein sequence database, which brings together experimental results, computed features and scientific conclusions."</ref>. Three iterations are performed using an expectation value cutoff of 10<sup>-3</sup>. | |||
# The sequences gathered with PSI-BLAST are then filtered (see below) using a scheme that attempts a balance between limiting the sequences to close homologues, and including distant sequences that do not share structure or function. | # The sequences gathered with PSI-BLAST are then filtered (see below) using a scheme that attempts a balance between limiting the sequences to close homologues, and including distant sequences that do not share structure or function. | ||
#The filtered sequence set is multiply aligned with [http://www.drive5.com/muscle/ MUSCLE] (a multiple sequence alignment algorithm that out-performs CLUSTALW). | #The filtered sequence set is multiply aligned with [http://www.drive5.com/muscle/ MUSCLE] (a multiple sequence alignment algorithm that out-performs CLUSTALW). | ||
| Line 58: | Line 59: | ||
#The normalized conservation scores are then ranked into nine levels from 1 (highly variable) to 9 (highly conserved). | #The normalized conservation scores are then ranked into nine levels from 1 (highly variable) to 9 (highly conserved). | ||
#Colors mapped to the nine conservation levels, from turquoise (1) to burgandy (9) are applied to the 3D protein structure visualized in [[FirstGlance in Jmol]]. | #Colors mapped to the nine conservation levels, from turquoise (1) to burgandy (9) are applied to the 3D protein structure visualized in [[FirstGlance in Jmol]]. | ||
Filtering of the sequences gathered for each protein chain is crucial to making the ConSurfDB results maximally informative. Filtering consists of the following steps. | |||
#Sequences with more than 95% sequence identity to the query sequence are discarded. | |||
#Sequences shorter than 60% of the query sequence are discarded. | |||
#Fragment sequences that overlap by over 10% are discarded. | |||
#Redundant sequences are removed using CD-HIT<ref>PMID: 16731699</ref>. | |||
#A maximum of 300 sequences meeting the above criteria is used (the 300 with the lowest expectation values, that is, most closely related to the query sequence). | |||
==The ConSurf Server== | ==The ConSurf Server== | ||