AlphaFold: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Angel Herraez (talk | contribs)
link to report at FEBS Network
Eric Martz (talk | contribs)
No edit summary
Line 3: Line 3:
*See [[Theoretical_models#2020:_CASP_14]] for more about the initial demonstration at CASP14, and the reactions to it.
*See [[Theoretical_models#2020:_CASP_14]] for more about the initial demonstration at CASP14, and the reactions to it.
*[[AlphaFold2_examples_from_CASP_14]] describes a detailed analysis of two of the CASP14 predictions.
*[[AlphaFold2_examples_from_CASP_14]] describes a detailed analysis of two of the CASP14 predictions.
==AlphaFold Database of Predictions==
In July, 2021, DeepMind made available over 300,000 structure predictions from amino acid sequences in their free [https://alphafold.ebi.ac.uk/ AlphaFold DB]<ref name="deepminddb">[https://deepmind.com/research/case-studies/alphafold#a_treasure_trove We’ve made AlphaFold predictions freely available to anyone in the scientific community] at DeepMind.com (date of release not specified, approximately July 2021).</ref><ref name="afdbebi">[https://www.ebi.ac.uk/pdbe/about/news/alphafold%E2%80%99s-protein-structure-predictions-now-available-explore AlphaFold’s protein structure predictions now available to explore] at the European Bioinformatics Institute, July 23, 2021.</ref><ref name="impacts">[https://www.embl.org/news/science/alphafold-potential-impacts/ Great expectations – the potential impacts of AlphaFold DB] at the European Bioinformatics Institute, July 22, 2021</ref><ref name="human">[https://www.embl.org/news/science/alphafold-database-launch/ DeepMind and EMBL release the most complete database of predicted 3D structures of human proteins] at the European Bioinformatics Institute, July 22, 2021.</ref>. These predictions include nearly all ~20,000 proteins in the human proteome, 36% with very high confidence, and another 22% with high confidence<ref name="human" /><ref name="human-nature">PMID: 34293799</ref>. Also included are ''E. coli'', fruit fly, mouse, zebrafish, malaria parasite and tuberculosis bacteria<ref name="human" />. Limitations of these predictions were enumerated<ref name="impacts" />, including:
* Inability to predict protein-protein or protein-DNA/RNA/ligand complexes. [[#RoseTTAFold]] claims to have made progress on this.
* Does not predict ligands, cofactors, metals, ions, glycosylation, etc.
* Does not deal with conformational dynamics.
* Does not predict [[Intrinsically Disordered Protein|intrinsically unstructured]] segments.
* Does not predict the folding pathway.
* Has not been trained to predict structural consequences of '''point mutations'''.
Nevertheless, these predictions have many potential benefits<ref name="impacts" />, including:
* Simplifying [[X-ray crystallography]] by enabling solution of the phase problem by molecular replacement using the predicted model.
* Assisting crystallographers in defining [[domain]] boundaries in order to crystallize domains when crystallization of full length proteins is problematic.
* Helping to interpret >5,000 [[cryo-EM]] maps previously deposited in the [[EMDB]] that could not be interpreted as atomic models, as well as helping to interpret lower resolution EM maps as atomic models.


==AlphaFold published July 2021==
==AlphaFold published July 2021==
Line 26: Line 41:


DeepMind has provided an [https://colab.research.google.com/github/deepmind/alphafold/blob/main/notebooks/AlphaFold.ipynb Alphafold Colab] that uses a "slightly simplified" version of AlphaFold version 2.0: "While accuracy will be near-identical to the full AlphaFold system on many targets, a small fraction have a large drop in accuracy due to the smaller MSA and lack of templates.". The AlphaFold Colab is '''free to use'''. The code is executed in a virtual machine private to your account, and data are stored on Google Drive. Nothing is installed on your computer; "everything happens in the cloud on Google Colab"<ref name="alphafoldcolab">[https://colab.research.google.com/github/deepmind/alphafold/blob/main/notebooks/AlphaFold.ipynb Alphafold Colab].</ref>
DeepMind has provided an [https://colab.research.google.com/github/deepmind/alphafold/blob/main/notebooks/AlphaFold.ipynb Alphafold Colab] that uses a "slightly simplified" version of AlphaFold version 2.0: "While accuracy will be near-identical to the full AlphaFold system on many targets, a small fraction have a large drop in accuracy due to the smaller MSA and lack of templates.". The AlphaFold Colab is '''free to use'''. The code is executed in a virtual machine private to your account, and data are stored on Google Drive. Nothing is installed on your computer; "everything happens in the cloud on Google Colab"<ref name="alphafoldcolab">[https://colab.research.google.com/github/deepmind/alphafold/blob/main/notebooks/AlphaFold.ipynb Alphafold Colab].</ref>
==AlphaFold Database of Predictions==
Also in July, 2021, DeepMind made available over 300,000 structure predictions from amino acid sequences in their free [https://alphafold.ebi.ac.uk/ AlphaFold DB]<ref name="deepminddb">[https://deepmind.com/research/case-studies/alphafold#a_treasure_trove We’ve made AlphaFold predictions freely available to anyone in the scientific community] at DeepMind.com (date of release not specified, approximately July 2021).</ref><ref name="afdbebi">[https://www.ebi.ac.uk/pdbe/about/news/alphafold%E2%80%99s-protein-structure-predictions-now-available-explore AlphaFold’s protein structure predictions now available to explore] at the European Bioinformatics Institute, July 23, 2021.</ref><ref name="impacts">[https://www.embl.org/news/science/alphafold-potential-impacts/ Great expectations – the potential impacts of AlphaFold DB] at the European Bioinformatics Institute, July 22, 2021</ref><ref name="human">[https://www.embl.org/news/science/alphafold-database-launch/ DeepMind and EMBL release the most complete database of predicted 3D structures of human proteins] at the European Bioinformatics Institute, July 22, 2021.</ref>. These predictions include nearly all ~20,000 proteins in the human proteome, 36% with very high confidence, and another 22% with high confidence<ref name="human" /><ref name="human-nature">PMID: 34293799</ref>. Also included are ''E. coli'', fruit fly, mouse, zebrafish, malaria parasite and tuberculosis bacteria<ref name="human" />. Limitations of these predictions were enumerated<ref name="impacts" />, including:
* Inability to predict protein-protein or protein-DNA/RNA/ligand complexes. [[#RoseTTAFold]] claims to have made progress on this.
* Does not predict ligands, cofactors, metals, ions, glycosylation, etc.
* Does not deal with conformational dynamics.
* Does not predict [[Intrinsically Disordered Protein|intrinsically unstructured]] segments.
* Does not predict the folding pathway.
* Has not been trained to predict structural consequences of '''point mutations'''.
Nevertheless, these predictions have many potential benefits<ref name="impacts" />, including:
* Simplifying [[X-ray crystallography]] by enabling solution of the phase problem by molecular replacement using the predicted model.
* Assisting crystallographers in defining [[domain]] boundaries in order to crystallize domains when crystallization of full length proteins is problematic.
* Helping to interpret >5,000 [[cryo-EM]] maps previously deposited in the [[EMDB]] that could not be interpreted as atomic models, as well as helping to interpret lower resolution EM maps as atomic models.


==References==
==References==