Unusual sequence numbering: Difference between revisions
From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs) No edit summary |
Eric Martz (talk | contribs) |
||
| (7 intermediate revisions by the same user not shown) | |||
| Line 1: | Line 1: | ||
The numbering of protein and nucleic acid sequences is arbitrary in structure files from the [[PDB|World Wide Protein Data Bank]] (PDB). That is, authors are free to number sequences as they wish. | The numbering of protein and nucleic acid sequences is arbitrary in structure files from the [[PDB|World Wide Protein Data Bank]] (PDB). That is, authors are free to number sequences as they wish. If you need to change the numbering in a published [[PDB file]], please see [[Renumbering PDB files]]. | ||
'''Straightforward numbering''' assigns 1 to the amino-terminal amino acid (or 5' nucleotide), and counts up sequentially and monotonically to the carboxy-terminal amino acid (or 3' nucleotide). An example is [http://firstglance.jmol.org/fg.htm?mol=1pgb 1pgb] ([[1pgb]]). The crystallized protein is numbered 1-56, despite it being a fragment of a [http://www.uniprot.org/uniprot/P06654#sequences 448-residue full length sequence] that begins (after adding an N-terminal Met) at full-length sequence number 228. | '''Straightforward numbering''' assigns 1 to the amino-terminal amino acid (or 5' nucleotide), and counts up sequentially and monotonically to the carboxy-terminal amino acid (or 3' nucleotide). An example is [http://firstglance.jmol.org/fg.htm?mol=1pgb 1pgb] ([[1pgb]]). The crystallized protein is numbered 1-56, despite it being a fragment of a [http://www.uniprot.org/uniprot/P06654#sequences 448-residue full length sequence] that begins (after adding an N-terminal Met) at full-length sequence number 228. | ||
| Line 33: | Line 33: | ||
===Insertion Codes In Reverse=== | ===Insertion Codes In Reverse=== | ||
Rarely, the insertion codes are in reverse alphabetical order. An example is [http://firstglance.jmol.org/fg.htm?mol=1ucy 1ucy] ([[1ucy]]). Chain L begins with nine amino acids all numbered 1. The insertion codes are in '''reverse-alphabetic order''': 1H, 1G, 1F, ... 1B, 1A, 1, 2, 3 .... In the same chain L are fourteen residues numbered 14. These insertion codes are in '''forward alphabetic order''': 13, 14, 14A, 14B, ... 14L, 14M, 15, 16 .... Chain | Rarely, the insertion codes are in reverse alphabetical order. An example is [http://firstglance.jmol.org/fg.htm?mol=1ucy 1ucy] ([[1ucy]]). Chain L begins with nine amino acids all numbered 1. The insertion codes are in '''reverse-alphabetic order''': 1H, 1G, 1F, ... 1B, 1A, 1, 2, 3 .... In the same chain L are fourteen residues numbered 14. These insertion codes are in '''forward alphabetic order''': 13, 14, 14A, 14B, ... 14L, 14M, 15, 16 .... Chain H has ten residues numbered 60, with forward-alphabetic insertion codes from A through I, and a few other shorter runs of insertion codes. | ||
==Gaps In Sequence Numbering== | ==Gaps In Sequence Numbering== | ||
| Line 44: | Line 44: | ||
===Missing Residues=== | ===Missing Residues=== | ||
[[Image:Sequence-missing-loop-2ace.png|frame|Excerpt from PDB file 2ace showing gap in sequence numbering due to a missing loop.]] | [[Image:Sequence-missing-loop-2ace.png|frame|Excerpt from PDB file 2ace showing gap in sequence numbering due to a missing loop.]] | ||
It is not uncommon for a surface loop of the crystallized protein to be disordered. Often such loops are [[Intrinsically Disordered Protein|intrinsically disordered]]. The disorder blurs the electron density map for that loop, and the loop residues are not given coordinates in the model: they are missing in the model. However, they were not missing in the crystallized protein. This causes a gap in the sequence numbers in the PDB file. An example is [http://firstglance.jmol.org/fg.htm?mol=2ace 2ace] ([[2ace]]). Residues 485-489 are missing in the 3D crystallographic model due to disorder in the crystal. Also missing are 3 N-terminal, and 2 C-terminal residues. FirstGlance in Jmol tabulates missing residues, and marks regions of the 3D model where residues are missing with "empty baskets". | It is not uncommon for a surface loop of the crystallized protein to be disordered. Often such loops are [[Intrinsically Disordered Protein|intrinsically disordered]]. The disorder blurs the electron density map for that loop, and the loop residues are not given coordinates in the model: they are [[Missing residues and incomplete sidechains|missing in the model]]. However, they were not missing in the crystallized protein. This causes a gap in the sequence numbers in the PDB file. An example is [http://firstglance.jmol.org/fg.htm?mol=2ace 2ace] ([[2ace]]). Residues 485-489 are missing in the 3D crystallographic model due to disorder in the crystal. Also missing are 3 N-terminal, and 2 C-terminal residues. FirstGlance in Jmol tabulates missing residues, and marks regions of the 3D model where residues are missing with "empty baskets". | ||
<table width=550><tr><td>[[Image:2ace-empty-basket.png|center]]</td><td>"Empty Basket": Closeup of the region of [[2ace]] where residues 485-489 are missing. In [[FirstGlance in Jmol]], empty baskets alert the user to missing residues. ("S-" labels residues with missing sidechain atoms.)</td></tr></table> | <table width=550><tr><td>[[Image:2ace-empty-basket.png|center]]</td><td>"Empty Basket": Closeup of the region of [[2ace]] where residues 485-489 are missing. In [[FirstGlance in Jmol]], empty baskets alert the user to missing residues. ("S-" labels residues with missing sidechain atoms.) | ||
<br><br> | |||
See also [[Missing residues and incomplete sidechains]].</td></tr></table> | |||
{{clear}} | {{clear}} | ||
| Line 62: | Line 64: | ||
== Notes == | == Notes == | ||
<references/> | <references/> | ||
==See Also== | |||
*[[Renumbering PDB files]] | |||
*[[Missing residues and incomplete sidechains]] | |||