Chains and Chain IDs: Difference between revisions
From Proteopedia
Jump to navigationJump to search
Eric Martz (talk | contribs) |
Eric Martz (talk | contribs) |
||
| (36 intermediate revisions by the same user not shown) | |||
| Line 12: | Line 12: | ||
Polypeptide ([[protein]]) chains are '''linear''', with rare exceptions where side-chains form [[protein crosslinks]] between two linear chains, | Polypeptide ([[protein]]) chains are '''linear''', with rare exceptions where side-chains form [[protein crosslinks]] between two linear chains, | ||
such as [[disulfide bonds]], or less commonly other types [[protein crosslinks]] | such as [[disulfide bonds]], or less commonly other types of [[protein crosslinks]], such as [[isopeptide bond]]s. | ||
Each protein chain has '''two ends''', an amino terminus (positively charged) and a carboxy terminus (negatively charged). The first residue in a protein chain becomes the amino terminus, with new amino acids being added at the carboxy terminus. The sequence of [[amino acids]] is specified by messenger RNA, which is a copy of the sequence of codons in the template strand of the DNA gene. The first residue in a nucleic acid chain becomes the 5' (phosphate) terminus, with new nucleotides being added at the 3' (hydroxy) terminus. | Each protein chain has '''two ends''', an amino terminus (positively charged) and a carboxy terminus (negatively charged). The first residue in a protein chain becomes the amino terminus, with new amino acids being added at the carboxy terminus. The sequence of [[amino acids]] is specified by messenger RNA, which is a copy of the sequence of codons in the template strand of the DNA gene. The first residue in a nucleic acid chain becomes the 5' (phosphate) terminus, with new nucleotides being added at the 3' (hydroxy) terminus. | ||
| Line 21: | Line 21: | ||
==Chain IDs== | ==Chain IDs== | ||
In the [[atomic coordinate file]]s maintained by the [[wwPDB]] ([[PDB files]]), each polymer chain is given an ID, or chain "name". In the legacy [[Atomic_coordinate_file#PDB_Data_Format|PDB data format]], chain IDs are a single letter or numeral (A-Z, a-z, 0-9), which limits the number of chains to 62. In the newer [[Atomic_coordinate_file#mmCIF_Data_Format|mmCIF data format]] (also called PDBx), chain IDs can be [https://mmcif.wwpdb.org/docs/large-pdbx-examples/index.html up to 4 letters or numerals], so the number of chains in a single structure | In the [[atomic coordinate file]]s maintained by the [[wwPDB]] ([[PDB files]]), each polymer chain is given an ID, or chain "name". In the legacy [[Atomic_coordinate_file#PDB_Data_Format|PDB data format]], chain IDs are a single letter or numeral (A-Z, a-z, 0-9), which limits the number of chains to 62. In the newer [[Atomic_coordinate_file#mmCIF_Data_Format|mmCIF data format]] (also called PDBx), chain IDs can be [https://mmcif.wwpdb.org/docs/large-pdbx-examples/index.html up to 4 letters or numerals], so the number of chains in a single structure has no practical limitation (>10 million chains/structure could be accommodated by 4-character chain IDs). | ||
===Oligosaccharide Chain IDs=== | ===Oligosaccharide Chain IDs=== | ||
The assignment of unique chain IDs to disaccharides and oligosaccharides began with the [https://www.wwpdb.org/documentation/remediation 2020 wwPDB Remediation of Carbohydrates]. Notably, '''monomeric''' nucleotides and amino acids (not part of a polymeric chain) and monosaccharides are assigned the chain ID of the '''nearest protein or nucleic acid''', while '''multimeric''' di- or oligo-nucleotides, and di- or oligosaccharides are given '''unique''' chain IDs. See item 2 below for the special case of dipeptides vs. tri- / oligo-peptides. | The assignment of unique chain IDs to disaccharides and oligosaccharides began with the [https://www.wwpdb.org/documentation/remediation 2020 wwPDB Remediation of Carbohydrates]. Notably, '''monomeric''' nucleotides and amino acids (not part of a polymeric chain) and monosaccharides are assigned the chain ID of the '''nearest protein or nucleic acid''', while '''multimeric''' di- or oligo-nucleotides, and di- or oligosaccharides are given '''unique''' chain IDs. See item 2 below for the special case of dipeptides vs. tri- / oligo-peptides. N-linked glycans are likely underrepresented in the PDB due to microheterogeneity and their flexibility<ref>PMID: 40645091</ref>. | ||
The procedure for assigning chain IDs is specified in the wwPDB Procedures section [https://www.wwpdb.org/documentation/procedure#toc_6 6. Chain ID assignment]. In | ===Chain ID Assignment Policies=== | ||
# When protein or nucleic acid is present, ligands and water bound to carbohydrate are '''never assigned the chain ID of that carbohydrate''', but are given the chain ID of the nearest protein/nucleic acid, even when it is >5 Å away (examples: [[7lkc|7LKC]], [[7dc4]], [[8g82]]). When the structure is carbohydrate without any protein or nucleic acid, only then are ligands and water given the chain ID of the nearest carbohydrate (examples: [[1c58]], [[2kqo]]). | |||
# Although dinucleotides and disaccharides are assigned unique chain IDs, dipeptides are not. Traditionally deemed ligands, | The procedure for assigning chain IDs is specified in the wwPDB Procedures section [https://www.wwpdb.org/documentation/procedure#toc_6 6. Chain ID assignment]. In April, 2025 that document needs two corrections in order to agree with actual wwPDB practice: | ||
# When protein or nucleic acid is present, ligands and water bound to carbohydrate are '''almost never<ref name="cacho">An exception is Ca318 in [[3gzt]], which is assigned chain ID X. Chain X is a disaccharide. In the asymmetric unit, Ca318 is 44 Å from the disaccharide, but only 26 Å from the nearest protein (chain Q). In Biomolecule 4, it is 2.6 Å from sidechain oxygens of Asp231 in chain B, and in the vicinity of two other chain B Asp.</ref> assigned the chain ID of that carbohydrate''', but are given the chain ID of the nearest protein/nucleic acid, even when it is >5 Å away (examples: [[7lkc|7LKC]], [[7dc4]], [[8g82]]). When the structure is carbohydrate without any protein or nucleic acid, only then are ligands and water given the chain ID of the nearest carbohydrate (examples: [[1c58]], [[2kqo]]). | |||
# Although dinucleotides and disaccharides are assigned unique chain IDs, '''dipeptides are not supposed to be assigned unique chain IDs.''' This policy is documented at the wwPDB in [https://www.wwpdb.org/documentation/procedure#toc_3 3. Polymer sequences and sequence database reference assignment]. '''However, this policy has not been followed consistently.''' A search at RCSB for ''Polymer Entity Sequence Length'' = 2 and ''Polymer Entity Type'' is Protein and Return ''Polymer Entities'' finds 59 hits (April, 2025), where dipeptides were given unique author-assigned chain IDS (in both PDB format and mmCIF format files). Traditionally deemed ligands, dipeptides are supposed to be assigned the same chain ID as the polymer chain to which they are bound. Dipeptide examples are [[2cyh]] and [[1dpp]]; tripeptide: [[4q1l|4q1L]]. Tripeptides and longer oligopeptides are supposed to be assigned unique chain IDs. When a dipeptide is not assigned a unique chain ID, it has no SEQRES and cannot be found by a search for polymer length at RCSB (April, 2025). | |||
===Author vs. wwPDB Chain IDs=== | ===Author vs. wwPDB Chain IDs=== | ||
| Line 34: | Line 36: | ||
CAUTION: Because the definitions of chain attributes are sometimes unclear in the [https://mmcif.wwpdb.org/pdbx-mmcif-search.html mmCIF Dictionaries], assertions in this section are the interpretations of a small sample of mmCIF files by [[User:Eric Martz]], and might contain errors. Please report any concerns or corrections to [[Image:Martz email.png|150px]]. | CAUTION: Because the definitions of chain attributes are sometimes unclear in the [https://mmcif.wwpdb.org/pdbx-mmcif-search.html mmCIF Dictionaries], assertions in this section are the interpretations of a small sample of mmCIF files by [[User:Eric Martz]], and might contain errors. Please report any concerns or corrections to [[Image:Martz email.png|150px]]. | ||
</td></tr></table> | </td></tr></table> | ||
An idiosyncracy of mmCIF | An idiosyncracy of '''mmCIF''' [[wwPDB]] files is that not only polymer chains of protein, nucleic acid, and carbohydrates, but all components in the structure model are assigned chain IDs, including ligands, metal ions, and water. '''Regardless of the chain IDs assigned by the authors''' of a structure model entry deposited in the wwPDB, the '''wwPDB assigns an additional set''' of (usually distinct) chain IDs. These PDBx chain IDs are present only in the mmCIF files, not in the PDB format files. In mmCIF files, there is no single place that lists all author-assigned or all wwPDB-assigned chain IDs. | ||
* <b>Author</b>-assigned chain IDs: | * <b>Author</b>-assigned chain IDs: | ||
| Line 222: | Line 224: | ||
</table> | </table> | ||
<font color="magenta">* Not compatible with PDB format</font> because of its 62-chain limit (A-Z, a-z, 0-9). Some author-assigned chain IDs must have at least 2 characters. | <font color="magenta">* Not compatible with PDB format</font> because of its 62-chain limit (A-Z, a-z, 0-9). Some author-assigned chain IDs must have at least 2 characters. | ||
===AlphaFold3 Chain IDs=== | |||
The [https://alphafoldserver.com AlphaFold Server], which in 2025 uses AlphaFold3<ref name="af3">PMID: 38718835</ref>, predicts complexes with multiple chains of protein and/or nucleic acid, plus a limited set of ligands, and metal ions, and a wide range of post-translational modifications of amino acids and chemical modifications of nucleotides (see [[How to predict structures with AlphaFold]]). Predicted models are available in '''mmCIF''' format only (not PDB format, although the mmCIF files can be [[Converting AlphaFold3 CIF to PDB|easily converted to PDB format]]). Consistent with the chain ID assignment policies of the wwPDB for mmCIF files, every entity is assigned a unique chain ID, including polymer chains, ligands, metal ions, and glycans including oligo- and monosaccharides. The following three examples are provided by the Server. | |||
<table class="wikitable"> | |||
<tr> | |||
<th> | |||
PDB ID | |||
</th> | |||
<th> | |||
Protein | |||
</th> | |||
<th> | |||
Author-assigned chain IDs in PDB & mmCIF files | |||
</th> | |||
<th> | |||
AlphaFold3 Server-assigned chain IDs | |||
</th> | |||
<th> | |||
Notes | |||
</th> | |||
</tr> | |||
<tr><!-- - - - - - - - - --> | |||
<td> | |||
[[7bbv]] | |||
</td> | |||
<td> | |||
Pectate lyase B | |||
<br> | |||
Biological Unit 1 | |||
</td> | |||
<td> | |||
1 ID total: | |||
:Protein, Zn++, Monomeric mannose glycoconjugates: <b>A</b>. | |||
</td> | |||
<td> | |||
8 IDs total: | |||
<br> | |||
:Protein: <b>A</b>. | |||
:1 Zn++: <b>B</b>. | |||
:6 Monomeric mannose glycoconjugates: <b>C,D,E,F,G,H</b>. | |||
</td> | |||
<td> | |||
Mannoses are conjugated to 5 threonines and 1 serine. | |||
</td> | |||
</tr> | |||
<tr><!-- - - - - - - - - --> | |||
<td> | |||
[[7rce]] | |||
</td> | |||
<td> | |||
Synthetic constructs | |||
</td> | |||
<td> | |||
3 IDs total: | |||
:Protein, Ca++, Na+: <b>A</b>. | |||
:DNA: <b>B, C</b>. | |||
</td> | |||
<td> | |||
7 IDs total: | |||
<br> | |||
:Protein: <b>A</b>. | |||
:3 Ca++: <b>B, C, D</b>. | |||
:1 Na+: <b>E</b>. | |||
:DNA: <b>F, G</b>. | |||
</td> | |||
<td> | |||
The 2.4 Å X-ray model has 85 amino acid sidechains missing distal atoms, including 50 charged residues, and is completely missing a small loop of 4 residues that includes one positive charge. These missing atoms are all present in the AlphaFold3-predicted model. | |||
</td> | |||
</tr> | |||
<tr><!-- - - - - - - - - --> | |||
<td> | |||
[[8aw3]] | |||
</td> | |||
<td> | |||
tRNA Deaminase | |||
</td> | |||
<td> | |||
3 IDs total: | |||
:tRNA: <b>1</b>. | |||
:Protein + 1 Zn++: <b>2</b>. | |||
:Protein + 1 Zn++: <b>3</b>. | |||
</td> | |||
<td> | |||
5 IDs total: | |||
<br> | |||
:Protein: <b>A, B</b>. | |||
:Zn++: <b>C, D</b>. | |||
:tRNA: <b>E</b>. | |||
</td> | |||
<td> | |||
</td> | |||
</tr> | |||
</table> | |||
==Notes== | |||
<references /> | |||
==See Also== | ==See Also== | ||