Amino acid composition: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
New page: The ''amino acid composition'' of a protein refers to the percentages of each amino acid in the sequence of that protein. The percentage, sometimes called the Mole percentage, is calculate...
 
Eric Martz (talk | contribs)
 
(37 intermediate revisions by the same user not shown)
Line 1: Line 1:
The ''amino acid composition'' of a protein refers to the percentages of each amino acid in the sequence of that protein. The percentage, sometimes called the Mole percentage, is calculated as the number of a given amino acid divided by the length of the protein chain, or total number of amino acids in the molecule.
The ''amino acid composition'' of a protein refers to the percentages of each amino acid in the sequence of that protein. The percentage, sometimes called the Mole percentage, is calculated for each of the [[amino acids|22 standard amino acids]] as the count of that amino acid divided by the total number of amino acids in the protein chain or molecule.


The strongest predictor of amino acid composition is the GC-content of the organism's genome.
==Example==
As an example, here is the amino acid composition of acetylcholinesterase of ''Torpedo californica'' (the Pacific electric ray), whose structure is [[2ace]]. The [https://www.uniprot.org/uniprot/P04058#sequences canonical isoform sequence] has length 586. In its mature form, a signal peptide is removed from the amino-terminus, and a pro-peptide is removed from the carboxy-terminus, leaving a mature length of 537, with this composition:
{| class="wikitable" style="margin-left: auto; margin-right: auto; border: none;width:710px;"
|-
|[[Image:Composition-by-pir-for-2ace.png|center]]
|-
|This composition bar graph was created by the Protein Information Resource's (PIR's) [https://proteininformationresource.org/pirwww/search/comp_mw.shtml Composition/Molecular Weight Calculator]. Protein sequences are easily obtained from [http://UniProt.Org UniProt.Org] or by viewing a PDB entry in [http://firstglance.jmol.org FirstGlance in Jmol] and clicking on Sequences. You may wish to align the genomic full-length sequence from UniProt with the experimentally crystallized sequence. Here are [http://firstglance.jmol.org/seqalign.htm instructions].
|}
 
==Average Compositions==
Average compositions have been calculated for large numbers of proteins from diverse taxa. These are tabulated in the downloadable spreadsheet [http://proteopedia.org/wiki/images/1/15/Amino-acid-composition.xlsx.zip amino-acid-composition.xlsx.zip]. It is reassuring to see the agreement between tabulations generated in 1993, 1998, and 2008 (citations are in the spreadsheet).
 
[[Image:Composition-caruso-200.png|770px|center]]
 
The above percentages were determined for several thousand sequences of diverse proteins of length 200 residues, with sequence identities below 50%<ref name="length" />. These data are included in the above-linked spreadsheet.
 
==Determinants of Amino Acid Composition==
 
'''GC-content''' of the organism's genome is the strongest genome-level determinant of amino acid composition.<ref name="tekala-genomes">PMID: 12384285</ref><ref name="linkers" /><ref name="habitats">PMID: 24204807</ref>.
 
Other, weaker influences are:
*'''Growth temperatures''' (mesophily/thermophily/hyperthermophily). Thermophiles have more glutamic acid (with reduction in glutamine), and more lysine and arginine<ref name="tekala-genomes" />. This likely relates to the larger number of [[salt bridges]] in proteins of thermophiles, believe to contribute to thermostability<ref name="saltbridges">PMID:21720566</ref>.
*'''Chain length'''. Proteins of thermophiles are, on average, shorter than those of mesophiles. Average lengths are 283 and 340, respectively<ref name="tekala-genomes" />. A study of ~550,000 proteins with lengths 50-200 amino acids<ref name="length">PMID:18780815</ref> concluded:
**Increased with length, reaching a plateau: Ala, Asp, Glu, Gly, Pro, Val; less increase for Gln and Thr.
**Decreased with length: Cys, Phe, His, Ile, Lys, Met, Asn, Ser.
**Leu and Tyr are highest in short and long chains, and less frequent in middle-sized proteins.
**Arg peaks in middle-sized proteins.
**Trp is constant at about 1.4% for lengths 75-200.
*'''Linkers vs. domains''': Linkers between domains have more polar residues, while compact domains have more hydrophobic residues<ref name="linkers">PMID: 29426365</ref>.
*'''Habitat''': The environment in which an organism lives has a minor effect on the average composition of its proteins<ref name="habitats" />.
*Compositional variability ranks archaea > baceteria > eukaryotes<ref name="linkers" />.
 
==Composition Calculators==
*Protein Information Resource's (PIR's) [https://proteininformationresource.org/pirwww/search/comp_mw.shtml Composition/Molecular Weight Calculator] makes a very useful bar graph (see example above) but does not provide a spreadsheet-ready table.
 
*EMBL-EBI's [https://www.ebi.ac.uk/Tools/seqstats/emboss_pepstats/ EMBOSS-PepStats] generates a table readily imported into a spreadsheet. The table has both 1-letter and 3-letter amino acid abbreviations, ''sorted by 1-letter codes''.
{| class="wikitable" style="margin-left: auto; margin-right: auto; border: none;width:70%;"
|'''Importing Composition Data Into Excel:''' Copy the data columns only, paste into a [[Help:plain text editors|plain text editor]] and save to a plain text file. In Excel, in an existing (possibly empty) spreadsheet, File, Import, Text. Check 3 delimiter options: Tab, Space, Treat consecutive delimiters as one. Proceed to import.
|}
 
*ExPASy's [https://web.expasy.org/cgi-bin/protparam/protparam ProtParam] generates a table readily imported into a spreadsheet. The table has both 1-letter and 3-letter amino acid abbreviations, ''sorted by 3-letter codes''. It also offers a CSV output, an alternative format understood by spreadsheets.
 
==References==
<references />