Mosaic Q
The Proteins Mosaic Q Pattern
Francisco Javier Lobo Cabrera, Ph.D (francisco.lobo6@gmail.com)
The Mosaic Q is a structural pattern observed in protein three-dimensional structures, in which amino acids of the same chemical type tend to cluster spatially in groups of specific size and geometry. Large-scale analysis of over 160,000 X-ray determined structures in the RCSB PDB suggests that this clustering can be described by a highly conserved empirical law (R² = 0.979) that depends primarily on protein size (number of residues).[1] Stochastic simulations prove that the experimental law that quantifies the pattern is dependent on a specific configuration of cluster size and geometry. In parallel, independent visual verification of this pattern is being carried out through the Proteins Mosaic Q Project, a citizen science initiative that enables distributed inspection of protein structures using standardized visualization protocols (project website,repository). The project has been registered and accepted in external platforms such as SciStarter[2], FAIRsharing[3],EU-Citizen.Science[4] and Open Science Framework[5], providing additional transparency and independent verification pathways.
What is the Mosaic Q Pattern? Alongside the well-known hydrophobic core, where hydrophobic residues tend to coalesce, residues of the same chemical category (polar, acidic, basic, or special) appear to assemble into short contiguous clusters, with an average length of around eight residues. This organization is referred to as the Mosaic Q model, in which clusters of approximately 8 residues of the same chemical type coexist within the protein structure alongside the hydrophobic core. At present, the model provides an empirical description of a statistical regularity observed in protein structures. The overall relationship between spatial clustering and protein size is illustrated schematically below. Stochastic simulations demonstrate that the shape of the experimental Q-n curve does depend on the number of residues per cluster and on cluster geometry. In this manner, it is shown that a good aproximation for the experimental results is a model where, in addition to the existence of a hydrophobic core, the rest of amino acids are grouped together according to their residue type (polar, basic, acidic, special) in clusters of approximately 8 residues. It must be noted, however, that the mosaic must be understood as a highly conserved pattern in protein structure, where exceptions may occur for different biological/evolutionary reasons, although the signal for this model is strong (R^2=0.978). Visual inspection of multiple PDB structures also point that the conclusions of the computational analysis are apparently valid. This is the basis for the validation mechanism by multiple independent observers, as discussed later, in the citizen-science aspect of this project. Indeed, the pattern can be visualized using a color scheme based on amino acid chemical type:
Amino acid chemical type classification follows established literature: acidic residues (Asp, Glu)[6], basic residues (Arg, His, Lys)[7], polar residues (Ser, Thr, Asn, Gln)[8], and hydrophobic residues (Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp)[9]. Special residues (Cys, Sec, Gly, Pro) are treated separately following Bak et al.[10], Arnér[11], Imamoto et al.[12] and Ho et al.[13] In all the examples, structures are rendered in spacefill Van der Waals representation, restricted to protein atoms, using the following Jmol script: select ala or val or ile or leu or met or phe or tyr or trp; color white; select ser or thr or asn or gln; color green; select asp or glu; color orange; select arg or his or lys; color blue; select cys or sec or gly or pro; color cyan; select all; spacefill vdw; restrict protein; ContentsQuantification: The Mosaic Q DescriptorThe degree of spatial clustering is quantified using the Mosaic Q descriptor. The position of each residue in space is defined by its α-carbon atom. Residues are classified into four chemical types: acidic (Asp, Glu), basic (Arg, His, Lys), polar (Ser, Thr, Asn, Gln) and hydrophobic (Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp). Special residues (Cys, Sec, Gly, Pro) are excluded from the standard Q calculation. Q is defined as the sum of squared inverse Euclidean distances between α-carbons of residues of the same chemical type. For each ordered pair of residues (i, j) of the same chemical type, the contribution is 1/d²(i,j), where d(i,j) is the Euclidean distance between the α-carbons of residues i and j, measured in Ångströms. This is summed over all ordered pairs within each chemical type, that is acidic (A), basic (B), polar (P) and hydrophobic (H): where for each type, i and j range independently over all residues of that type (i ≠ j). Note that each pair is counted twice in this formulation, consistent with the original definition. An alternative descriptor, Q_alt, also includes special residues: where S denotes special residues (Cys, Sec, Gly, Pro), and i and j range over all residues of that type (i ≠ j). Higher Q values indicate stronger spatial clustering, although their interpretation should be considered in relation to protein size. The empirical law found relates Q linearly to the number of residues (n) with very limited variance (R² = 0.978), indicating that proteins of similar size tend to display similar Q values, largely independent of fold, conformation or organism of origin: An analogous empirical law holds for Q_alt: The Q descriptor is implemented in the protein-mosaic-q PyPI package, also registered in bio.tools[14]. Given a PDB or CIF file, it can be computed as follows: mosaicq protein.pdb mosaicq protein.cif --metric Q_alt where the
Source code is available at GitHub. Exploring the pattern's presence in well-known proteinsAlthough the statistical analysis suggests that the Mosaic Q pattern is broadly conserved across protein structures, it is instructive to visualize it in some well-known examples spanning different sizes, folds and biological functions. Accordingly, this Proteopedia entry enables interactive 3D exploration of the Mosaic Q pattern across representative protein structures, enabling the Proteopedia community, with its deep expertise in protein structural biology, to assess and contribute to the project's effort. Click on any protein name below to load it in the viewer:
Independent Verification Through Citizen ScienceThe Proteins Mosaic Q Project is an open citizen science initiative designed to facilitate independent verification of the observed pattern through reproducible visualization workflows. The project is registered in multiple external platforms and open science registries, including: Participants can generate independent visual evidence by applying a standardized representation to protein structures from the RCSB PDB. This approach does not replace statistical analysis, but provides a complementary layer of transparency and reproducibility through distributed inspection of individual cases. References
| ||||||||||||