Intrinsically Disordered Protein: Difference between revisions
From Proteopedia
Jump to navigationJump to search
m fix linebreak |
No edit summary |
||
| Line 2: | Line 2: | ||
It has long been taught that proteins must be properly folded in order to perform their functions. This paradigm derives from work by Christian B. Anfinsen and coworkers. In the 1960's, they showed that RNAse, when denatured so that 99% of its enzymatic activity was lost, could regain enzymatic activity within seconds when the denaturing agent was removed under proper conditions<ref>For the sake of brevity, this description is oversimplified. RNAse needed to be reduced to break disulfide bonds, as well as using 8 M urea, for denaturation. Oxidation without the denaturant then left an inactive enzyme because the disulfide bonds formed randomly, precluding proper folding except very slowly (many hours). Only when protein disulfide isomerase was added did the re-folding occur at a physiological rate (about a minute). The fact that RNAse could thus be trapped in an inactive conformation under physiological conditions contributed to the insights developed by Anfinsen and his team. Proteins lacking disulfides renatured in seconds. For details, see [http://nobelprize.org/nobel_prizes/chemistry/laureates/1972/anfinsen-lecexture.html Anfinsen's Nobel Lecture.]</ref><ref>A similar observation was made around the same time by then graduate student Lisa Steiner in the lab of [[Richards, Frederic M.|Fred Richards]] at Yale University. Neither Richards nor advisor Joseph Fruton thought the observation interesting enough to publish. It was an answer to a question not yet asked. This story is recounted by David Eisenberg, see the next citation.</ref><ref>PMID: 29958112</ref>. They concluded that the amino acid sequence is sufficient for a protein to fold into its functional, lowest energy conformation. This work won the [[Nobel_Prizes_for_3D_Molecular_Structure|1972 Nobel Prize]], and was subsequently confirmed and extended by many researchers. | It has long been taught that proteins must be properly folded in order to perform their functions. This paradigm derives from work by Christian B. Anfinsen and coworkers. In the 1960's, they showed that RNAse, when denatured so that 99% of its enzymatic activity was lost, could regain enzymatic activity within seconds when the denaturing agent was removed under proper conditions<ref>For the sake of brevity, this description is oversimplified. RNAse needed to be reduced to break disulfide bonds, as well as using 8 M urea, for denaturation. Oxidation without the denaturant then left an inactive enzyme because the disulfide bonds formed randomly, precluding proper folding except very slowly (many hours). Only when protein disulfide isomerase was added did the re-folding occur at a physiological rate (about a minute). The fact that RNAse could thus be trapped in an inactive conformation under physiological conditions contributed to the insights developed by Anfinsen and his team. Proteins lacking disulfides renatured in seconds. For details, see [http://nobelprize.org/nobel_prizes/chemistry/laureates/1972/anfinsen-lecexture.html Anfinsen's Nobel Lecture.]</ref><ref>A similar observation was made around the same time by then graduate student Lisa Steiner in the lab of [[Richards, Frederic M.|Fred Richards]] at Yale University. Neither Richards nor advisor Joseph Fruton thought the observation interesting enough to publish. It was an answer to a question not yet asked. This story is recounted by David Eisenberg, see the next citation.</ref><ref>PMID: 29958112</ref>. They concluded that the amino acid sequence is sufficient for a protein to fold into its functional, lowest energy conformation. This work won the [[Nobel_Prizes_for_3D_Molecular_Structure|1972 Nobel Prize]], and was subsequently confirmed and extended by many researchers. | ||
Beginning around 2000, it was recognized that not all proteins function in a folded state<ref>PMID: 10550212</ref><ref>PMID: 11381529</ref><ref>PMID: 11784292</ref><ref name="tompa2002">PMID: 12368089</ref><ref>Summary of the previous paper (Tompa, 2002): The disorder of intrinsically | Beginning around 2000, it was recognized that not all proteins function in a folded state<ref>PMID: 10550212</ref><ref>PMID: 11381529</ref><ref>PMID: 11784292</ref><ref name="tompa2002">PMID: 12368089</ref><ref>Summary of the previous paper (Tompa, 2002): The disorder of intrinsically disordered proteins (IDP's) is crucial to their functions. They may adopt defined but extended structures when bound to cognate ligands. Their amino acid compositions are less hydrophobic than those of soluble proteins. They lack hydrophobic cores, and hence do not become insoluble when heated. About 40% of eukaryotic proteins have at least one long (>50 residues) disordered region. Roughly 10% of proteins in various genomes have been predicted to be fully disordered. Presently over 100 IDP's have been identified; none are enzymes. Obviously, IDP's are greatly underrepresented in the Protein Data Bank, although there are a few cases of an IDP bound to a folded (intrinsically structured) protein. Here, Tompa suggests five functional categories for intrinsically unstructured proteins and domains: entropic chains (bristles to ensure spacing, springs, flexible spacers/linkers), effectors (inhibitors and disassemblers), scavengers, assemblers, and display sites. (Summary by Eric Martz.)</ref><ref>PMID: 18952168</ref><ref>PMID: 15284216</ref>. Some proteins must be unfolded or disordered in order to perform their functions, and others fold only in complex with target structures<ref name="gunasekaran2003">PMID: 12575995</ref><ref>Summary of the previous paper (Gunasekaran ''et al.'', 2003): Argues that proteins involved in extensive protein-protein interactions can function effectively despite having their structure depend upon such interactions, so that as monomers they are natively disordered. Dispensing with the structural framework (scaffold) needed to maintain a stable fold in the monomer increases efficiency by reducing size. This may account for the large percentage (roughly half) of all proteins that are predicted to be natively disordered. (Summary by Eric Martz.)</ref><ref>PMID: 15738986</ref>. These are termed '''intrinsically disordered protein (IDP), intrinsically unstructured protein (IDP), or natively unfolded protein'''. | ||
By some estimates, about 10% of all proteins are fully disordered, and about 40% of eukaryotic proteins have at least one long (>50 amino acids) disordered loop<ref name="tompa2002" />. Such sequences, under physiological conditions ''in vitro'', display physicochemical characteristics resembling those of random coils. They possess little or no ordered structure, having instead an extended conformation with high intra-molecular flexibility, lacking any tightly packed core. | By some estimates, about 10% of all proteins are fully disordered, and about 40% of eukaryotic proteins have at least one long (>50 amino acids) disordered loop<ref name="tompa2002" />. Such sequences, under physiological conditions ''in vitro'', display physicochemical characteristics resembling those of random coils. They possess little or no ordered structure, having instead an extended conformation with high intra-molecular flexibility, lacking any tightly packed core. | ||
| Line 26: | Line 26: | ||
Many [[X-ray crystallography|crystallographic]] structures have missing loops -- that is, ranges of amino acids with no [[atomic coordinate file|atomic coordinates]] in the model. These "gaps" in the model are often thought to be artifacts of inadvertant disorder in the crystal. In some cases, these gaps may be alerting us to the presence of intrinsically disordered loops in an otherwise folded protein<ref name="IDSG" />. Such gaps are the basis for the [[#Protein disorder predictors|DISOPRED2 disorder prediction server]]. [[FirstGlance in Jmol]] offers [[Temperature_value#Missing_Residues|one method for locating and visualizaing such gaps]]. | Many [[X-ray crystallography|crystallographic]] structures have missing loops -- that is, ranges of amino acids with no [[atomic coordinate file|atomic coordinates]] in the model. These "gaps" in the model are often thought to be artifacts of inadvertant disorder in the crystal. In some cases, these gaps may be alerting us to the presence of intrinsically disordered loops in an otherwise folded protein<ref name="IDSG" />. Such gaps are the basis for the [[#Protein disorder predictors|DISOPRED2 disorder prediction server]]. [[FirstGlance in Jmol]] offers [[Temperature_value#Missing_Residues|one method for locating and visualizaing such gaps]]. | ||
Despite the existence of compelling evidence for | Despite the existence of compelling evidence for IDPs and intrinsically disordered loops beginning in 1990<ref name="struhl1990" /><ref>PMID: 2236048</ref><ref>For the ''unstructured domain'' interpretation of early work by Pontius and Berg, see the 2004 review by Tompa and Csermley, PMID: 15284216</ref>, many current textbooks of biochemistry and even some monographs on protein structure fail to mention intrinsic disorder and its importance for protein function<ref>PMID: 18831774</ref><ref>Martz, E. Book review of <i>Introduction to protein science—architecture, function, and genomics: Lesk, Arthur M.</i>. <i>Biochem. Mol. Biol. Educ.</i> 33:144-5 (2006). [http://dx.doi.org/10.1002/bmb.2005.494033022442 DOI: 10.1002/bmb.2005.494033022442]</ref>. In 2011, Chouard provided a readable and informative overview of IDPs and how some of them function<ref>PMID: 21390105</ref>. | ||
__TOC__ | __TOC__ | ||
== Examples of | == Examples of IDPs == | ||
Examples cover a wide variety of cellular systems and it has been predicted that eukaryotes have more | Examples cover a wide variety of cellular systems and it has been predicted that eukaryotes have more IDPs than other kingdoms <ref>PMID: 11700597</ref>. Of course, there are no [[PDB codes]] for fully disordered proteins in isolation. However, there are some crystallographic results for IDP that undergo [[#Many IDPs undergo disorder-order transition|disorder-order transition]] when they complex with another folded protein domain, such as [[1jsu]], [[1g3j]], and [[1oct]]<ref name="tompa2002" />. Other examples are at [[Globular_Proteins]]. See further information about [[1jsu]] and other cases [[#Many IDPs undergo disorder-order transition|below]]. | ||
IDPs play roles in '''processes''' such as: | |||
* Cell signaling and cell cycle regulation, e.g. cyclin dependent kinase inhibitor p21Waf1/Cip1/Sdi1<ref>PMID: 8876165</ref> | * Cell signaling and cell cycle regulation, e.g. cyclin dependent kinase inhibitor p21Waf1/Cip1/Sdi1<ref>PMID: 8876165</ref> | ||
| Line 61: | Line 61: | ||
<table align='right'><tr valign='top'><td> | <table align='right'><tr valign='top'><td> | ||
[[Image:FoldIndex.jpeg | thumb | FoldIndex<ref name="foldindex">PMID: 15955783</ref> output for three protein sequences (a) Cat-Muscle Pyruvate Kinase (b) The human p53 tumor suppressor protein (c) Chicken gizzard caldesmon; green is folded and red is unfolded]]</td><td>[[Image:DisorderAA.jpg | thumb | Content of order-promoting and disorder-promoting amino acids in the ''Drosophila'' proteome (black) and in the cytoplasmic domain of gliotactin that was shown to be | [[Image:FoldIndex.jpeg | thumb | FoldIndex<ref name="foldindex">PMID: 15955783</ref> output for three protein sequences (a) Cat-Muscle Pyruvate Kinase (b) The human p53 tumor suppressor protein (c) Chicken gizzard caldesmon; green is folded and red is unfolded]]</td><td>[[Image:DisorderAA.jpg | thumb | Content of order-promoting and disorder-promoting amino acids in the ''Drosophila'' proteome (black) and in the cytoplasmic domain of gliotactin that was shown to be IDP (gray) <ref>PMID: 14579366</ref>]] | ||
</td></tr></table> | </td></tr></table> | ||
{{Clear}} | {{Clear}} | ||
Led by the assumption that “since amino acid sequence determines 3-D structure, amino acid sequence should also determine lack of 3-D structure” <ref name='Dunker2001'>PMID: 11533628</ref> specific sequence features shared by | Led by the assumption that “since amino acid sequence determines 3-D structure, amino acid sequence should also determine lack of 3-D structure” <ref name='Dunker2001'>PMID: 11533628</ref> specific sequence features shared by IDPs have been evaluated and algorithms for their identification formulated. | ||
The low hydrophobicity and high [[net charge]] of natively unfolded proteins result in a difference in amino acid composition between them and natively folded proteins <ref>PMID: 11093259</ref>. | The low hydrophobicity and high [[net charge]] of natively unfolded proteins result in a difference in amino acid composition between them and natively folded proteins <ref>PMID: 11093259</ref>. | ||
| Line 86: | Line 86: | ||
* [http://bioinf.cs.ucl.ac.uk/disopred/ DISOPRED2] (Jones Group, University College London, UK). "DISOPRED2 was trained on a set of around 750 non-redundant sequences with high resolution X-ray structures. Disorder was identified with those residues that appear in the sequence records but with coordinates missing from the electron density map. This is an imperfect means for identifying disordered residues as missing co-ordinates can also arise as an artifact of the crystalization process. False assignment of order can also occur as a result of stabilizing interactions by ligands or other macromolecules in the complex. However, this is the simplest means for defining disorder in the absence of further experimental investigation of the protein." (Quoted from the DISOPRED2 website.) | * [http://bioinf.cs.ucl.ac.uk/disopred/ DISOPRED2] (Jones Group, University College London, UK). "DISOPRED2 was trained on a set of around 750 non-redundant sequences with high resolution X-ray structures. Disorder was identified with those residues that appear in the sequence records but with coordinates missing from the electron density map. This is an imperfect means for identifying disordered residues as missing co-ordinates can also arise as an artifact of the crystalization process. False assignment of order can also occur as a result of stabilizing interactions by ligands or other macromolecules in the complex. However, this is the simplest means for defining disorder in the absence of further experimental investigation of the protein." (Quoted from the DISOPRED2 website.) | ||
* [http://bip.weizmann.ac.il/fldbin/findex/ FoldIndex]<ref name="foldindex" /> (Sussman Group, Weizmann Institute, Rehovot, Israel). FoldIndex makes predictions based on the observation that | * [http://bip.weizmann.ac.il/fldbin/findex/ FoldIndex]<ref name="foldindex" /> (Sussman Group, Weizmann Institute, Rehovot, Israel). FoldIndex makes predictions based on the observation that IDPs occupy the low hydrophobicity/ high net-charge portion of charge-hydrophobicity phase space. (See Figure above.) | ||
* [http://iupred.enzim.hu/ IUPred] (Dosztányi, Csizmók, Tompa and Simon: Budapest, Hungary). "IUPred recognized intrinsically unstructured regions from the amino acid sequence based on the estimated pairwise energy content. The underlying assumption is that globular proteins are composed of amino acids which have the potential to form a large number of favorable interactions, whereas intrinsically | * [http://iupred.enzim.hu/ IUPred] (Dosztányi, Csizmók, Tompa and Simon: Budapest, Hungary). "IUPred recognized intrinsically unstructured regions from the amino acid sequence based on the estimated pairwise energy content. The underlying assumption is that globular proteins are composed of amino acids which have the potential to form a large number of favorable interactions, whereas intrinsically disorered proteins (IDPs) adopt no stable structure because their amino acid composition does not allow sufficient favorable interactions to form." (Quoted from the IUPred website.) | ||
* [http://www.pondr.com/ PONDR] (Dunker Group, Indiana University and Molecular Kinetics, Inc., Indianapolis IN USA; Obradovic Group, Temple Univ., Philadelphia PA USA). "PONDR® functions from primary sequence data alone. The predictors are feedforward neural networks that use sequence information from windows of generally 21 amino acids. Attributes, such as the fractional composition of particular amino acids or hydropathy, are calculated over this window, and these values are used as inputs for the predictor. The neural network, which has been trained on a specific set of ordered and disordered sequences, then outputs a value for the central amino acid in the window. The predictions are then smoothed over a sliding window of 9 amino acids. If a residue value exceeds a threshold of 0.5 (the threshold used for training) the residue is considered disordered." (Quoted from the PONDR website.) | * [http://www.pondr.com/ PONDR] (Dunker Group, Indiana University and Molecular Kinetics, Inc., Indianapolis IN USA; Obradovic Group, Temple Univ., Philadelphia PA USA). "PONDR® functions from primary sequence data alone. The predictors are feedforward neural networks that use sequence information from windows of generally 21 amino acids. Attributes, such as the fractional composition of particular amino acids or hydropathy, are calculated over this window, and these values are used as inputs for the predictor. The neural network, which has been trained on a specific set of ordered and disordered sequences, then outputs a value for the central amino acid in the window. The predictions are then smoothed over a sliding window of 9 amino acids. If a residue value exceeds a threshold of 0.5 (the threshold used for training) the residue is considered disordered." (Quoted from the PONDR website.) | ||
| Line 113: | Line 113: | ||
== Biological implications of | == Biological implications of IDPs == | ||
It was proposed that the unfolded nature of the | It was proposed that the unfolded nature of the IDPs provides them with advantages in recognition and binding. Although their large hydrodynamic dimensions slow down diffusion, their size provides a large target for initial molecular collisions, and the lack of rigid binding pockets permits multiple approach orientations for a binding partner, which may increase the probability of productive interactions <ref>PMID: 12065587</ref><ref name='Dunker2001'/>. In addition, IDPs allow molecular plasticity by adopting more than one conformation and binding diversity by binding to several proteins and thus many of the known hub proteins are IDPs. IDPs rapid turnover in the cell allow their tight regulation as many times needed in cell signaling and cell cycle. | ||
== Evolution of | == Evolution of IDPs == | ||
In p53, the folded DNA-binding domain is conserved, while the intrinsically disordered regions display a higher rate of mutations<ref>PMID: 23352836</ref>. | In p53, the folded DNA-binding domain is conserved, while the intrinsically disordered regions display a higher rate of mutations<ref>PMID: 23352836</ref>. | ||
== Many | == Many IDPs undergo disorder-order transition == | ||
Binding of natural ligands such as a variety of small molecules, substrates, cofactors, other proteins, nucleic acids or membranes may induce unstructured proteins to adopt stable structures bound to the partner, or even a secondary structure bound to the partner. In addition to the cases detailed below, other examples include [[1g3j]], [[1oct]]<ref name="tompa2002" />, and the [[Lac repressor]]. | Binding of natural ligands such as a variety of small molecules, substrates, cofactors, other proteins, nucleic acids or membranes may induce unstructured proteins to adopt stable structures bound to the partner, or even a secondary structure bound to the partner. In addition to the cases detailed below, other examples include [[1g3j]], [[1oct]]<ref name="tompa2002" />, and the [[Lac repressor]]. | ||
Some | Some IDP sequences are able to bind to multiple partners that have <25% sequence identity, and in some cases even different folds<ref name="one-to-many">PMID: 23233352</ref>. For example, the C-terminal portion of p53 is known to bind to four different protein partners each with different folds<ref name="one-to-many" />; and the N-terminus of histone H3 binds to nine different protein partners with distinct folds<ref name="one-to-many" />. | ||
=== The human p27<sup>Kip1</sup> kinase inhibitory domain <ref>PMID: 8684460</ref> === | === The human p27<sup>Kip1</sup> kinase inhibitory domain <ref>PMID: 8684460</ref> === | ||
| Line 132: | Line 132: | ||
The cyclin-dependent kinases (CDKs) have a central role in coordinating the eukaryotic cell division cycle. CDKs are controlled through several different processes involving the binding of activating cyclin subunits. Complexes of cyclins with CDKs play a central role in the control of the eukaryotic cell cycle. These complexes are inhibited by other proteins termed in general cyclin-CDK inhibitors (CKIs). One example of CKIs is p27<sup>Kip1</sup>. p27<sup>Kip1</sup> is an | The cyclin-dependent kinases (CDKs) have a central role in coordinating the eukaryotic cell division cycle. CDKs are controlled through several different processes involving the binding of activating cyclin subunits. Complexes of cyclins with CDKs play a central role in the control of the eukaryotic cell cycle. These complexes are inhibited by other proteins termed in general cyclin-CDK inhibitors (CKIs). One example of CKIs is p27<sup>Kip1</sup>. p27<sup>Kip1</sup> is an IDP and it binds to phosphorylated <scene name='User:Tzviya_Zeev-Ben-Mordehai/Sandbox_1/Complex/2'>cyclin/CDK complex</scene> in <scene name='User:Tzviya_Zeev-Ben-Mordehai/Sandbox_1/Extended/2'>an extended conformation</scene> interacting with both <scene name='User:Tzviya_Zeev-Ben-Mordehai/Sandbox_1/Cyca/2'>cyclin A</scene> and <scene name='User:Tzviya_Zeev-Ben-Mordehai/Sandbox_1/Cdk2/3'>CDK2</scene> ([[1jsu]]). On cyclin A, it binds in a groove formed by conserved cyclin box residues. On CDK2, it binds and rearranges the amino-terminal lobe and also inserts into the catalytic cleft, mimicking ATP. [[http://www.proteopedia.org/wiki/index.php/1jsu]] | ||
{{Clear}} | {{Clear}} | ||
| Line 140: | Line 140: | ||
The yeast transcriptional activator GCN4 belongs to a large family of eukaryotic transcription factors including Fos, Jun and CREB. All family members have a [[DNA]] recognition motif consists of a coiled-coil dimerization element, the leucine-zipper, and an adjoining basic region, which mediates DNA binding. This basic region is largely unstructured in the absence of DNA, addition of DNA containing a GCN4 binding site induce the transition of this region from unstructured to α-helical<ref>PMID: 12381856</ref>. | The yeast transcriptional activator GCN4 belongs to a large family of eukaryotic transcription factors including Fos, Jun and CREB. All family members have a [[DNA]] recognition motif consists of a coiled-coil dimerization element, the leucine-zipper, and an adjoining basic region, which mediates DNA binding. This basic region is largely unstructured in the absence of DNA, addition of DNA containing a GCN4 binding site induce the transition of this region from unstructured to α-helical<ref>PMID: 12381856</ref>. | ||
{{Clear}} | {{Clear}} | ||
== Practical Implications of | == Practical Implications of IDPs == | ||
There is evidence that large intrinsically unstructured regions interfere with crystallization<ref name="IDSG">PMID: 23232152</ref>. Oldfield ''et al.'', 2013<ref name="IDSG" />, concluded: | There is evidence that large intrinsically unstructured regions interfere with crystallization<ref name="IDSG">PMID: 23232152</ref>. Oldfield ''et al.'', 2013<ref name="IDSG" />, concluded: | ||
<blockquote> | <blockquote> | ||