Amino acid composition: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
Eric Martz (talk | contribs)
 
(17 intermediate revisions by the same user not shown)
Line 1: Line 1:
The ''amino acid composition'' of a protein refers to the percentages of each amino acid in the sequence of that protein. The percentage, sometimes called the Mole percentage, is calculated for each of the [[amino acids|22 standard amino acids]] as the number of a that amino acid divided by the total number of amino acids in the protein chain or molecule.
The ''amino acid composition'' of a protein refers to the percentages of each amino acid in the sequence of that protein. The percentage, sometimes called the Mole percentage, is calculated for each of the [[amino acids|22 standard amino acids]] as the count of that amino acid divided by the total number of amino acids in the protein chain or molecule.


==Example==
==Example==
As an example, here is the amino acid composition of acetylcholinesterase of ''Torpedo californica'' (the Pacific electric ray), whose structure is [[2ace]]. The [https://www.uniprot.org/uniprot/P04058#sequences canonical isoform sequence] has length 586. In its mature form, a signal peptide is removed from the amino-terminus, and a pro-peptide is removed from the carboxy-terminus, leaving a mature length of 537, with this composition:
As an example, here is the amino acid composition of acetylcholinesterase of ''Torpedo californica'' (the Pacific electric ray), whose structure is [[2ace]]. The [https://www.uniprot.org/uniprot/P04058#sequences canonical isoform sequence] has length 586. In its mature form, a signal peptide is removed from the amino-terminus, and a pro-peptide is removed from the carboxy-terminus, leaving a mature length of 537, with this composition:
{| class="wikitable" style="margin-left: auto; margin-right: auto; border: none;"
{| class="wikitable" style="margin-left: auto; margin-right: auto; border: none;width:710px;"
|-
|-
|[[Image:Composition-by-pir-for-2ace.png]]
|[[Image:Composition-by-pir-for-2ace.png|center]]
|-
|-
|This composition bar graph was created by the Protein Information Resource's (PIR's) [https://proteininformationresource.org/pirwww/search/comp_mw.shtml Composition/Molecular Weight Calculator].
|This composition bar graph was created by the Protein Information Resource's (PIR's) [https://proteininformationresource.org/pirwww/search/comp_mw.shtml Composition/Molecular Weight Calculator]. Protein sequences are easily obtained from [http://UniProt.Org UniProt.Org] or by viewing a PDB entry in [http://firstglance.jmol.org FirstGlance in Jmol] and clicking on Sequences. You may wish to align the genomic full-length sequence from UniProt with the experimentally crystallized sequence. Here are [http://firstglance.jmol.org/seqalign.htm instructions].
|}
|}
==Average Compositions==
Average compositions have been calculated for large numbers of proteins from diverse taxa. These are tabulated in the downloadable spreadsheet [http://proteopedia.org/wiki/images/1/15/Amino-acid-composition.xlsx.zip amino-acid-composition.xlsx.zip]. It is reassuring to see the agreement between tabulations generated in 1993, 1998, and 2008 (citations are in the spreadsheet).
[[Image:Composition-caruso-200.png|770px|center]]
The above percentages were determined for several thousand sequences of diverse proteins of length 200 residues, with sequence identities below 50%<ref name="length" />. These data are included in the above-linked spreadsheet.


==Determinants of Amino Acid Composition==
==Determinants of Amino Acid Composition==
Line 23: Line 30:
**Trp is constant at about 1.4% for lengths 75-200.
**Trp is constant at about 1.4% for lengths 75-200.
*'''Linkers vs. domains''': Linkers between domains have more polar residues, while compact domains have more hydrophobic residues<ref name="linkers">PMID: 29426365</ref>.
*'''Linkers vs. domains''': Linkers between domains have more polar residues, while compact domains have more hydrophobic residues<ref name="linkers">PMID: 29426365</ref>.
*'''Habitat''': The environment in which an organism lives has a minor effect on the average composition of its proteins<ref name="habitat" />.
*'''Habitat''': The environment in which an organism lives has a minor effect on the average composition of its proteins<ref name="habitats" />.
*Compositional variability ranks archaea > baceteria > eukaryotes<ref name="linkers" />.
*Compositional variability ranks archaea > baceteria > eukaryotes<ref name="linkers" />.
==Composition Calculators==
*Protein Information Resource's (PIR's) [https://proteininformationresource.org/pirwww/search/comp_mw.shtml Composition/Molecular Weight Calculator] makes a very useful bar graph (see example above) but does not provide a spreadsheet-ready table.
*EMBL-EBI's [https://www.ebi.ac.uk/Tools/seqstats/emboss_pepstats/ EMBOSS-PepStats] generates a table readily imported into a spreadsheet. The table has both 1-letter and 3-letter amino acid abbreviations, ''sorted by 1-letter codes''.
{| class="wikitable" style="margin-left: auto; margin-right: auto; border: none;width:70%;"
|'''Importing Composition Data Into Excel:''' Copy the data columns only, paste into a [[Help:plain text editors|plain text editor]] and save to a plain text file. In Excel, in an existing (possibly empty) spreadsheet, File, Import, Text. Check 3 delimiter options: Tab, Space, Treat consecutive delimiters as one. Proceed to import.
|}
*ExPASy's [https://web.expasy.org/cgi-bin/protparam/protparam ProtParam] generates a table readily imported into a spreadsheet. The table has both 1-letter and 3-letter amino acid abbreviations, ''sorted by 3-letter codes''. It also offers a CSV output, an alternative format understood by spreadsheets.


==References==
==References==
<references />
<references />