AlphaFold2 examples from CASP 14: Difference between revisions
Eric Martz (talk | contribs) No edit summary |
Eric Martz (talk | contribs) No edit summary |
||
| (8 intermediate revisions by the same user not shown) | |||
| Line 1: | Line 1: | ||
Prediction of protein structures from amino acid sequences, [[theoretical modeling]], has been extremely challenging. In 2020, breakthrough success was achieved by AlphaFold2<ref name="af2">PMID:31942072</ref>, a project of [http://deepmind.com DeepMind]. '''For an overview of this breakthrough''', documented by the bi-annual prediction competition [[Theoretical_models#CASP|CASP]], please see [[Theoretical_models#2020:_CASP_14|2020: CASP 14]]. Below are illustrated two examples of predictions from that competition. | Prediction of protein structures from amino acid sequences, [[theoretical modeling]], has been extremely challenging. In 2020, breakthrough success was achieved by AlphaFold2<ref name="af2">PMID:31942072</ref>, a project of [http://deepmind.com DeepMind]. '''For an overview of this breakthrough''', documented by the bi-annual prediction competition [[Theoretical_models#CASP|CASP]], please see [[Theoretical_models#2020:_CASP_14|2020: CASP 14]]. Below are illustrated two examples of predictions from that competition. | ||
| Line 86: | Line 84: | ||
===T1037 contains several known fold fragments=== | ===T1037 contains several known fold fragments=== | ||
(Dali | The X-ray structure of T1037 (404 residues from 6vr4) was submitted to Dali<ref name="dali2020" /> in March, 2021. Among the ~1,000 hits with Z ≥ 2.0, there were 152 with lengths ≥ 400 residues, and 224 with lengths ≥ 300, long enough that a superposition with the majority of T1037 would not be precluded. Among all hits, the largest number of aligned residues was 140/404 (35%) with RMSD 11.7 Å. The second largest was 127/404 (31%), RMSD 7.7 Å. Thus, no single structure in the PDB superposed with more than 35% of T1037. | ||
However, several of the Dali hits superposed with non-overlapping core fragments of [[6vr4]]<ref name="lholm">These non-overlapping core fragments were kindly pointed out by Liisa Holm, March, 2021.</ref>: | |||
*[[2j7n]] chain A, RNA-dependent RNA polymerase | |||
**length 934, aligned residues '''115, RMSD 4.3 Å''', Z=4.0, structural alignment 9 %id. | |||
*[[4ncj]] chain A, DNA double-strand break repair RAD50 ATPase | |||
**length 311, aligned residues '''109, RMSD 4.7 Å''', Z=3.4, structural alignment 11 %id. | |||
*[[5vfk]] chain A, Uncharacterized protein | |||
**length 146, aligned residues '''61, RMSD 7.8 Å''', Z=3.3, structural alignment 11 %id. | |||
Liisa Holm<ref name="dali2020" /><ref name="holmquote">Quoted with permission from Liisa Holm, March, 2021.</ref> stated: "T1037 has a homologous template in the PDB. The parent structure of T1037, phage RNA polymerase (6vr4, 2166 amino acids), is homologous to the RNAi polymerase from Neurospora crassa (2j7n chain A, 934 amino acids)<ref name="6vr4" />. Dali aligns them over 564 residues with an RMSD of 4.8 A. 115 residues of the common core are in the T1037 substructure. Several long insertions in T1037/6vr4 relative to 2j7n (chain A) form subdomains, which point outwards from the common core. Similar massive adaptation of the common core is seen, for example, in the glucosyltransferase 1 family<ref>PMID: 7729407</ref>." | |||
The [https://fatcat.godziklab.org/ FATCAT Server] reported that in order to superpose 150 residues (37% of 404) of T1037 with the closest structure in the PDB, 3 twists at hinges were required, after which an RMSD of 3.1 Å was achieved. For a 200-residue superposition (50% of 404), the best results after 3 twists had an RMSD of 5.4 Å. | The [https://fatcat.godziklab.org/ FATCAT Server] reported that in order to superpose 150 residues (37% of 404) of T1037 with the closest structure in the PDB, 3 twists at hinges were required, after which an RMSD of 3.1 Å was achieved. For a 200-residue superposition (50% of 404), the best results after 3 twists had an RMSD of 5.4 Å. | ||
| Line 184: | Line 192: | ||
For comparison, CASP 14 reported GDT_TS 86.96 for the AlphaFold2 prediction, while the AS2TS server calculated GDT_TS 85.87 vs. 7jx6 chain A, and 88.32 vs. 7JTL chain A. (These results were corrected for 90/92 and 91/92 residues, respectively.) Thus, there appears to be some unidentified minor discrepancy between the GDT_TS calculations of CASP-14 vs. the method detailed at [[Calculating GDT_TS]]. | For comparison, CASP 14 reported GDT_TS 86.96 for the AlphaFold2 prediction, while the AS2TS server calculated GDT_TS 85.87 vs. 7jx6 chain A, and 88.32 vs. 7JTL chain A. (These results were corrected for 90/92 and 91/92 residues, respectively.) Thus, there appears to be some unidentified minor discrepancy between the GDT_TS calculations of CASP-14 vs. the method detailed at [[Calculating GDT_TS]]. | ||
==See Also== | |||
*[[AlphaFold/Index]], a list of pages in Proteopedia about Alphafold. | |||
==References & Notes== | ==References & Notes== | ||
<references /> | <references /> | ||
Latest revision as of 00:08, 29 September 2023
Prediction of protein structures from amino acid sequences, homology modeling, has been extremely challenging. In 2020, breakthrough success was achieved by AlphaFold2[1], a project of DeepMind. For an overview of this breakthrough, documented by the bi-annual prediction competition GDT_TS scores, please see 6vr4. Below are illustrated two examples of predictions from that competition.
| |||||||||||
ORF8 Sidechain Accuracy
AlphaFold2's predictions for sidechain positions seem fairly good, while sidechain positions in the 2nd best prediction seem poor. This conclusion is based on three types of observations:
- Table I gives RMSD values for all atoms, which is one indication of sidechain accuracy.
- Prediction of SARS-CoV-2 protein ORF8 and homology modeling.
- Visualization of the distributions of charges on the surfaces.
Salt Bridges and Cation-Pi Interactions
- AlphaFold2's prediction was correct for 4/5 interactions, with one incorrect interaction.
- AlphaFold2's prediction was correct for one of two salt bridges, and predicted no incorrect salt bridges.
- AlphaFold2's prediction was correct for three of three cation-pi interactions, but predicted one incorrect interaction.
- The 2nd best prediction was correct for 1/5 interactions, with 2 incorrect interactions.
- The 2nd best prediction was correct for one of two salt bridges, but predicted two incorrect salt bridges.
- The 2nd best prediction failed to predict any of the three cation-pi interactions, predicting zero interactions.
| 7JX6 | 7JTL | AlphaFold2 | 2nd Best |
|---|---|---|---|
| R101:D112 (AB) | R101:D113 (AB) | R86:D98 | R86:D98 |
| R115:D119 (AB) | R115:D119 (AB) | – | R100:E4 |
| K44:E59 (AB) | K44:E59 (AB) | K29:E44 | – |
| – | – | – | K78:E77 |
- Bridges in the same row are identical (except for red residues). Subtract 15 from the sequence numbers in the X-ray structures for the equivalent sequence numbers in the predictions.
- Black: Shortest sidechain nitrogen to sidechain oxygen distance ≤4.0 Å.
- Gray: Shortest sidechain nitrogen to sidechain oxygen distance 4.4 to 4.8 Å.
- –: Shortest sidechain nitrogen to sidechain oxygen distance 6 to 16 Å.
- (AB): The two chains in each X-ray model.
- Italics: erroneous prediction.
| 7JX6 | 7JTL | AlphaFold2 | 2nd Best |
|---|---|---|---|
| R101:Y46+Y108 (AB) | R101:Y46+Y108 (AB) | R86:Y31+Y96 | – |
| K44:F108 (B) | K44:F108 (AB) | K29:F93 | – |
| – | – | K79:F105 | – |
- All interactions listed are deemed energetically significant by the CaPTURE Server.
- Interactions in the same row are identical. Subtract 15 from the sequence numbers in the X-ray structures for the equivalent sequence numbers in the predictions.
- Italics: erroneous prediction.
- The 2nd best prediction has no cation-pi interactions.
- (AB): The two chains in each X-ray model.
Visualization of Surface Charge Distributions
The distributions of surface charges are in good agreement between AlphaFold2's prediction and the two crystal structures, which agree with each other. The distribution in the 2nd best prediction has several discrepancies with the other three models.
GDT_TS Calculations
GDT_TS values for predictions are taken from CASP 14 results. The reference structure for the CASP 14 GDT_TS values was 92 alpha carbons of 7JTL[2], since the CASP 14 target had only 92 residues[2].
GDT_TS values for 7JTL and 5A2F vs. 7JX6 chain A were calculated using the AS2TS server of Adam Zemla[3]. See instructions for empirical models. GDT_TS values were corrected for 92 residues (not 104) because the CASP 14 target had only 92 residues[2].
For comparison, CASP 14 reported GDT_TS 86.96 for the AlphaFold2 prediction, while the AS2TS server calculated GDT_TS 85.87 vs. 7jx6 chain A, and 88.32 vs. 7JTL chain A. (These results were corrected for 90/92 and 91/92 residues, respectively.) Thus, there appears to be some unidentified minor discrepancy between the GDT_TS calculations of CASP-14 vs. the method detailed at 7jtl.
See Also
- 7jx6, a list of pages in Proteopedia about Alphafold.
References & Notes
- ↑ Senior AW, Evans R, Jumper J, Kirkpatrick J, Sifre L, Green T, Qin C, Zidek A, Nelson AWR, Bridgland A, Penedones H, Petersen S, Simonyan K, Crossan S, Kohli P, Jones DT, Silver D, Kavukcuoglu K, Hassabis D. Improved protein structure prediction using potentials from deep learning. Nature. 2020 Jan;577(7792):706-710. doi: 10.1038/s41586-019-1923-7. Epub 2020 Jan, 15. PMID:31942072 doi:https://dx.doi.org/10.1038/s41586-019-1923-7
- ↑ 2.0 2.1 2.2 Cite error: Invalid
<ref>tag; no text was provided for refs namedcasp14domains - ↑ Zemla A. LGA: A method for finding 3D similarities in protein structures. Nucleic Acids Res. 2003 Jul 1;31(13):3370-4. doi: 10.1093/nar/gkg571. PMID:12824330 doi:https://dx.doi.org/10.1093/nar/gkg571