Unusual sequence numbering: Difference between revisions

From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs)
No edit summary
Eric Martz (talk | contribs)
No edit summary
Line 7: Line 7:
==Numbering Does Not Start With One==
==Numbering Does Not Start With One==
===N-Terminal Residues Missing Coordinates===
===N-Terminal Residues Missing Coordinates===
Probably the most common reason that the first residue with coordinates is not numbered 1 is because the N-terminal (or 5'-terminal) residues are missing coordinates due to crystallographic disorder (fuzzy electron density map). An example is [http://firstglance.jmol.org/fgij/fg.htm?1d66 1d66] ([[1d66]]). The first 7 residues of chain A are missing, so the first residue with coordinates is numbered 8. 1-7 were present in the crystallized protein, but could not be resolved in the electron density map.
Probably the most common reason that the first residue with coordinates is not numbered 1 is because the N-terminal (or 5'-terminal) residues are missing coordinates due to crystallographic disorder (fuzzy electron density map). An example is [http://firstglance.jmol.org/fg.htm?mol=1d66 1d66] ([[1d66]]). The first 7 residues of chain A are missing, so the first residue with coordinates is numbered 8. 1-7 were present in the crystallized protein, but could not be resolved in the electron density map.


===N-Terminal Residues Deleted From Protein===
===N-Terminal Residues Deleted From Protein===
Another common reason that sequence numbering does not start with 1 is because a range of N-terminal residues were deleted from the cloned and expressed protein used in the experiment. An example is chain A in [http://firstglance.jmol.org/fgij/fg.htm?1b07 1b07] ([[1b07]]). This 65 amino acid chain starts with Gly132-Ser133 that are not part of the gene sequence. Next comes '''Ala134''', and its sequence number (and the numbering of the remainder of the chain) '''matches the numbering''' of the [http://www.uniprot.org/uniprot/Q64010#sequences gene-encoded protein], full length 304 amino acids.
Another common reason that sequence numbering does not start with 1 is because a range of N-terminal residues were deleted from the cloned and expressed protein used in the experiment. An example is chain A in [http://firstglance.jmol.org/fg.htm?mol=1b07 1b07] ([[1b07]]). This 65 amino acid chain starts with Gly132-Ser133 that are not part of the gene sequence. Next comes '''Ala134''', and its sequence number (and the numbering of the remainder of the chain) '''matches the numbering''' of the [http://www.uniprot.org/uniprot/Q64010#sequences gene-encoded protein], full length 304 amino acids.


Authors do not always use the full-length sequence numbering when the structure of a fragment is reported. As mentioned above, in [http://firstglance.jmol.org/fgij/fg.htm?1pgb 1pgb] ([[1pgb]]), the crystallized protein is numbered 1-56. This despite it being a fragment of a [http://www.uniprot.org/uniprot/P06654#sequences 448-residue full length sequence] that begins (after adding an N-terminal Met) at full-length sequence number 228.
Authors do not always use the full-length sequence numbering when the structure of a fragment is reported. As mentioned above, in [http://firstglance.jmol.org/fg.htm?mol=1pgb 1pgb] ([[1pgb]]), the crystallized protein is numbered 1-56. This despite it being a fragment of a [http://www.uniprot.org/uniprot/P06654#sequences 448-residue full length sequence] that begins (after adding an N-terminal Met) at full-length sequence number 228.


===Starts With Zero Or Negative Numbers===
===Starts With Zero Or Negative Numbers===
'''Zero.''' Sometimes the initial sequence number is zero. An example is [http://firstglance.jmol.org/fgij/fg.htm?1bxw 1bxw] ([[1bxw]]). The first 21 residues of the [http://www.uniprot.org/uniprot/P0A910#sequences genomic sequence] are a signal sequence. The crystallized protein was engineered to start at residue 22 of the genomic sequence, which is Ala1 of the mature protein. A Met was engineered onto the N-terminus presumably to assist with expression. It was numbered Met0. (The crystallized protein ends at 178, but the length of the genomic sequence of the mature protein is 346 - 21 = 325.)
'''Zero.''' Sometimes the initial sequence number is zero. An example is [http://firstglance.jmol.org/fg.htm?mol=1bxw 1bxw] ([[1bxw]]). The first 21 residues of the [http://www.uniprot.org/uniprot/P0A910#sequences genomic sequence] are a signal sequence. The crystallized protein was engineered to start at residue 22 of the genomic sequence, which is Ala1 of the mature protein. A Met was engineered onto the N-terminus presumably to assist with expression. It was numbered Met0. (The crystallized protein ends at 178, but the length of the genomic sequence of the mature protein is 346 - 21 = 325.)


'''Negative.''' Sometimes the initial sequence number is negative. This is usually done when residues were engineered onto the N-terminus. The transition from -1 to 1 may or may not include a residue numbered zero. An example is [http://firstglance.jmol.org/fgij/fg.htm?1d5t 1d5t] ([[1d5t]]). The N-terminal Met of the [http://www.uniprot.org/uniprot/P21856#sequences genomic sequence] is numbered 1. But a di-histidine tag was engineered onto the N-terminus: His -2, His -1, Met 1. In this case, there is no residue numbered zero. The C-terminal residue is Phe431, but the length of the genomic sequence is 447. The C-terminal 16 residues of the genomic sequence were not present in the crystallized protein. In this model, no residues are missing due to crystallographic disorder.
'''Negative.''' Sometimes the initial sequence number is negative. This is usually done when residues were engineered onto the N-terminus. The transition from -1 to 1 may or may not include a residue numbered zero. An example is [http://firstglance.jmol.org/fg.htm?mol=1d5t 1d5t] ([[1d5t]]). The N-terminal Met of the [http://www.uniprot.org/uniprot/P21856#sequences genomic sequence] is numbered 1. But a di-histidine tag was engineered onto the N-terminus: His -2, His -1, Met 1. In this case, there is no residue numbered zero. The C-terminal residue is Phe431, but the length of the genomic sequence is 447. The C-terminal 16 residues of the genomic sequence were not present in the crystallized protein. In this model, no residues are missing due to crystallographic disorder.


==Multiple Residues with the Same Number==
==Multiple Residues with the Same Number==
===Insertion Codes===
===Insertion Codes===
[[Image:Sequence-insertion-codes-1igy.png|frame|Excerpt from PDB file 1igy showing insertion codes.]]
[[Image:Sequence-insertion-codes-1igy.png|frame|Excerpt from PDB file 1igy showing insertion codes.]]
Sometimes the residues of a protein are numbered according to a different ''reference sequence''. When there are insertions relative to the reference sequence, the additional residues may all be given the same sequence number, but marked with alphabetic insertion codes. This is frequently done in antibodies, where the reference sequence is the germline sequence, but the antibody has been somatically mutated, especially in complementarity-determining region (CDR) 3. An example is [http://firstglance.jmol.org/fgij/fg.htm?1igy 1igy] ([[1igy]]). Four residues in chain B all have sequence number 82. They are distinguished by insertion codes: 82, 82A, 82B, 82C. At right is this part of the PDB file.
Sometimes the residues of a protein are numbered according to a different ''reference sequence''. When there are insertions relative to the reference sequence, the additional residues may all be given the same sequence number, but marked with alphabetic insertion codes. This is frequently done in antibodies, where the reference sequence is the germline sequence, but the antibody has been somatically mutated, especially in complementarity-determining region (CDR) 3. An example is [http://firstglance.jmol.org/fg.htm?mol=1igy 1igy] ([[1igy]]). Four residues in chain B all have sequence number 82. They are distinguished by insertion codes: 82, 82A, 82B, 82C. At right is this part of the PDB file.


===Insertion Codes In Reverse===
===Insertion Codes In Reverse===