Unusual sequence numbering: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
No edit summary
Eric Martz (talk | contribs)
No edit summary
Line 25: Line 25:


===Insertion Codes In Reverse===
===Insertion Codes In Reverse===
Rarely, the insertion codes are in reverse alphabetical order. An example is [http://firstglance.jmol.org/fgij/fg.htm?1ucy 1ucy] ([[1ucy]]). Chain L begins with nine amino acids all numbered 1. The insertion codes are in '''reverse-alphabetic order''': 1H, 1G, 1F, ... 1B, 1A, 1, 2, 3 .... In the same chain L are fourteen residues numbered 14. These insertion codes are in '''forward alphabetic order''': 13, 14, 14A, 14B, ... 14L, 14M, 15, 16 .... Chain L also has ten residues numbered 60, with forward-alphabetic insertion codes from A through I, and a few other shorter runs of insertion codes.
Rarely, the insertion codes are in reverse alphabetical order. An example is [http://firstglance.jmol.org/fg.htm?mol=1ucy 1ucy] ([[1ucy]]). Chain L begins with nine amino acids all numbered 1. The insertion codes are in '''reverse-alphabetic order''': 1H, 1G, 1F, ... 1B, 1A, 1, 2, 3 .... In the same chain L are fourteen residues numbered 14. These insertion codes are in '''forward alphabetic order''': 13, 14, 14A, 14B, ... 14L, 14M, 15, 16 .... Chain L also has ten residues numbered 60, with forward-alphabetic insertion codes from A through I, and a few other shorter runs of insertion codes.


==Gaps In Sequence Numbering==
==Gaps In Sequence Numbering==


===Skipping Sequence Numbers===
===Skipping Sequence Numbers===
Sometimes a range of sequence numbers is skipped when numbering a continuous protein chain. There is no gap in the protein chain, but merely a discontinuity in the numbering of the chain. In the case of antibody [http://firstglance.jmol.org/fgij/fg.htm?1igy 1igy] ([[1igy]]), the sequence is numbered according to the Kabat scheme, relative to a reference sequence. Chain B begins with 1 and ends with 474 but contains only 444 residues (none are missing coordinates due to disorder). In chain B, residue 97 is followed by residue 100, '''skipping numbers 98-99'''. Only the numbers are skipped. No residues are missing. Residue 97 is peptide-bonded to residue 100. There are four residues 100, with insertion codes H, I, J, K. Residue 157 is followed by residue 162, '''skipping numbers 158-161'''. '''Also skipped are sequence numbers 170, 181-182, 197, 201, 207, 224-225, 233-234, 293-294, 297-298, 315-316, 356, 362, 376, 380, 403-404, 409, 412-413, 429, 431-432''', and probably more.
Sometimes a range of sequence numbers is skipped when numbering a continuous protein chain. There is no gap in the protein chain, but merely a discontinuity in the numbering of the chain. In the case of antibody [http://firstglance.jmol.org/fg.htm?mol=1igy 1igy] ([[1igy]]), the sequence is numbered according to the Kabat scheme, relative to a reference sequence. Chain B begins with 1 and ends with 474 but contains only 444 residues (none are missing coordinates due to disorder). In chain B, residue 97 is followed by residue 100, '''skipping numbers 98-99'''. Only the numbers are skipped. No residues are missing. Residue 97 is peptide-bonded to residue 100. There are four residues 100, with insertion codes H, I, J, K. Residue 157 is followed by residue 162, '''skipping numbers 158-161'''. '''Also skipped are sequence numbers 170, 181-182, 197, 201, 207, 224-225, 233-234, 293-294, 297-298, 315-316, 356, 362, 376, 380, 403-404, 409, 412-413, 429, 431-432''', and probably more.


{{Clear}}
{{Clear}}
===Missing Residues===
===Missing Residues===
[[Image:Sequence-missing-loop-2ace.png|frame|Excerpt from PDB file 2ace showing gap in sequence numbering due to a missing loop.]]
[[Image:Sequence-missing-loop-2ace.png|frame|Excerpt from PDB file 2ace showing gap in sequence numbering due to a missing loop.]]
It is not uncommon for a surface loop of the crystallized protein to be disordered. Often such loops are [[Intrinsically Disordered Protein|intrinsically disordered]]. The disorder blurs the electron density map for that loop, and the loop residues are not given coordinates in the model: they are missing in the model. However, they were not missing in the crystallized protein. This causes a gap in the sequence numbers in the PDB file. An example is [http://firstglance.jmol.org/fgij/fg.htm?2ace 2ace] ([[2ace]]). Residues 485-489 are missing in the 3D crystallographic model due to disorder in the crystal. Also missing are 3 N-terminal, and 2 C-terminal residues.  FirstGlance in Jmol tabulates missing residues, and marks regions of the 3D model where residues are missing with "empty baskets".
It is not uncommon for a surface loop of the crystallized protein to be disordered. Often such loops are [[Intrinsically Disordered Protein|intrinsically disordered]]. The disorder blurs the electron density map for that loop, and the loop residues are not given coordinates in the model: they are missing in the model. However, they were not missing in the crystallized protein. This causes a gap in the sequence numbers in the PDB file. An example is [http://firstglance.jmol.org/fg.htm?mol=2ace 2ace] ([[2ace]]). Residues 485-489 are missing in the 3D crystallographic model due to disorder in the crystal. Also missing are 3 N-terminal, and 2 C-terminal residues.  FirstGlance in Jmol tabulates missing residues, and marks regions of the 3D model where residues are missing with "empty baskets".


{{clear}}
{{clear}}
Line 41: Line 41:
==Not Monotonic==
==Not Monotonic==
[[Image:Sequence-not-monotonic-4zwj.png|frame|Excerpt from PDB file 4zwj showing non-monotonic sequence numbering in chain A.]]
[[Image:Sequence-not-monotonic-4zwj.png|frame|Excerpt from PDB file 4zwj showing non-monotonic sequence numbering in chain A.]]
Rarely, sequence numbers do not increase monotonically from N to C terminus. An example<ref>Thanks to Rachel Kramer Green of [[RCSB]] for this example.</ref> is [http://firstglance.jmol.org/fgij/fg.htm?4zwj 4zwj] ([[4zwj]]). In this chimeric protein, chain A is numbered 1002-1161 continuing 1-326 continuing 2012-2361. That is, there are sudden jumps in numbering of consecutive amino acids: 1161 to 1, and 326 to 2012. At right is an excerpt from the ATOM records of the [[PDB file]] for 4zwj chain A.
Rarely, sequence numbers do not increase monotonically from N to C terminus. An example<ref>Thanks to Rachel Kramer Green of [[RCSB]] for this example.</ref> is [http://firstglance.jmol.org/fg.htm?mol=4zwj 4zwj] ([[4zwj]]). In this chimeric protein, chain A is numbered 1002-1161 continuing 1-326 continuing 2012-2361. That is, there are sudden jumps in numbering of consecutive amino acids: 1161 to 1, and 326 to 2012. At right is an excerpt from the ATOM records of the [[PDB file]] for 4zwj chain A.


== References ==
== References ==
<references/>
<references/>