Protein Structure PAE Matrix: Meaning and Applications
Published on August 30, 2026

In the field of protein structure prediction, the emergence of deep learning methods such as the AlphaFold series (AlphaFold2/3) has brought revolutionary breakthroughs. Researchers can now rapidly obtain massive three-dimensional predicted structures of proteins. However, predicted structures are not experimental structures—how to evaluate the reliability of prediction results and distinguish reliable structural regions from random predictions is a core prerequisite for structural biology research. The protein PAE matrix (Predicted Aligned Error matrix) is the core quantitative tool for addressing this issue.
Abstract
Structure prediction tools such as the AlphaFold series have been widely applied in protein science research, but the reliability assessment of predicted structures remains a critical step constraining the quality of downstream analyses. The protein structure PAE matrix (Predicted Aligned Error Matrix) is the core quantitative tool for addressing this problem. The PAE matrix quantifies, in the form of an N×N matrix, the predicted positional error between any two amino acid residues in a protein. Unlike pLDDT, which is a per-residue local confidence metric, the PAE matrix focuses on the global spatial reliability of residue pairs. PAE values are reported in Ångströms (Å), with lower values indicating higher prediction accuracy and confidence for the relative positions of residues. By interpreting PAE heatmaps, researchers can precisely determine the folding reliability of protein domains, define domain boundaries, evaluate the overall conformational confidence of multi-domain proteins, and assess the authenticity of protein complex interactions. This article systematically reviews the definition of the PAE matrix, methods for visual interpretation, and its core application scenarios and usage guidelines in structural biology and protein engineering.
I. What Is the Protein Structure PAE Matrix?

Difference between PAE matrix and pLDDT metric
1.1 Basic Definition of PAE
The PAE matrix stands for Predicted Aligned Error Matrix. For a protein containing N amino acid residues, the PAE is an N×N two-dimensional square matrix, where the element PAE[i][j] is defined as: after globally aligning the predicted structure to the true native structure at residue j, the expected positional error of the Cα atom of residue i in three-dimensional space (in Å).
In simpler terms, PAE[i][j] answers the core question: using the spatial position of residue j as an anchor point and correcting for overall structural deviation, what is the average deviation between the predicted position and the true position of residue i? The core of this definition is relative alignment error, which has no direct correspondence with the physical distance between residues—a low PAE does not mean that residues are physically close, nor does a high PAE mean they are far apart. PAE matrices can be obtained directly from the AlphaFold Database, or generated locally by running AlphaFold or ColabFold.
1.2 Core Differences Between the PAE Matrix and pLDDT
In the AlphaFold series output system, pLDDT and PAE are two complementary and irreplaceable confidence systems:
pLDDT (Predicted Local Distance Difference Test): A per-residue local metric, with a score range of 0–100, that evaluates the folding accuracy of the local peptide chain surrounding a single residue. It cannot reflect the spatial arrangement between residues or between domains. pLDDT above 90 indicates high confidence, 70–90 indicates medium confidence, and below 70 indicates low confidence.
PAE matrix: A residue-pair global metric that specifically evaluates whether the relative spatial positions between residue pairs and between domains are truly reliable.
In short: pLDDT reflects the local folding quality of individual residues, while PAE reflects the global spatial reliability between residue pairs. High-quality local folding does not guarantee reliable global arrangement—the two must be used together.
1.3 PAE Value Range and Classification Criteria
The upper limit of PAE values natively output by AlphaFold is 31.75 Å, and the values are inversely correlated with prediction confidence. The classification criteria are as follows: PAE below 5 Å indicates a high-confidence region, where the relative positions of residue pairs are highly reliable and suitable for detailed structural analysis; PAE between 5 and 15 Å indicates a medium-confidence region, where the overall conformation has reference value but should be interpreted with caution; PAE above 15 Å indicates a low-confidence region, where the spatial arrangement of residues is essentially unreliable and should not be used as a basis for functional conclusions. The above thresholds are general reference standards; in specific studies, they should be applied flexibly according to the scenario—detailed analyses should strictly adhere to the 5 Å threshold, while large-scale screening may relax the criteria appropriately.
II. How to Read the Protein PAE Matrix Heatmap

Interpretation rules for PAE heatmap
The primary visualization format for the PAE matrix is a residue heatmap, with both axes representing residue numbers arranged by sequence order, and the color gradient corresponding to PAE values (dark green = low PAE = high confidence; light green/white = high PAE = low confidence).
2.1 Main Diagonal: A Zero-Error Baseline with No Biological Significance
The main diagonal of the heatmap is consistently dark green (PAE ≈ 0), which is determined by definition—the diagonal element PAE[i][i] represents alignment of a residue with itself, with a theoretical error of zero, and contains no structural information. This diagonal region should be ignored when interpreting the heatmap.
2.2 Blocks Adjacent to the Diagonal: Independent Rigid Domain Regions
Continuous, well-defined dark green square blocks appearing on both sides of the main diagonal correspond to independent rigid domains or stably folded segments of the protein. All residue pairs within these blocks exhibit low PAE values, indicating that the relative positions of residues are highly fixed and stably folded, representing the most reliable structural units in the prediction. The boundaries of these square blocks typically correspond to the natural boundaries of domains or flexible linker regions.
2.3 Off-Diagonal Regions: The Core Criterion for Inter-Domain Packing Confidence
The off-diagonal regions are the most scientifically valuable areas of the PAE heatmap for interpretation: dark green blocks spanning different regions (low PAE) indicate that the spatial packing and relative orientation between different domains are predicted to be truly reliable; light green/white blocks spanning regions (high PAE) suggest that the relative positions between domains are uncertain—while AlphaFold can accurately predict the folding of individual domains, it cannot determine the overall spatial assembly mode of multi-domain proteins.
2.4 Asymmetry of the PAE Matrix and Its Biological Significance
The PAE matrix is not strictly symmetric, meaning PAE[i][j] ≠ PAE[j][i]. This asymmetry is more pronounced in flexible loop regions, terminal disordered regions, and variable conformational linker regions. This feature directly reflects the dynamic uncertainty and anisotropic flexibility of regional conformations, and may indirectly suggest the presence of dynamic conformational changes in the protein. However, such inferences require validation through experimental evidence such as NMR or molecular dynamics simulations.
III. Typical Application Scenarios of the Protein PAE Matrix

Typical application scenarios of PAE matrix
3.1 Evaluating the Overall Conformational Confidence of Multi-Domain Proteins
Multi-domain proteins are the predominant form of eukaryotic proteins, and it is impossible to distinguish between "true packing" and "random prediction" based solely on the three-dimensional structural model. If the inter-domain PAE values are high, it indicates that the spatial arrangement has no predictive confidence and is the result of random fitting by the algorithm, and thus cannot be used for analyzing domain-domain interactions, functional site cooperativity, or other core questions. The PAE matrix can identify conformationally reliable models at the global level.
3.2 Precise Identification of Protein Domain Boundaries
Clustering algorithms based on the PAE matrix (such as Leiden clustering) enable data-driven domain partitioning: residue clusters with extremely low PAE values and high spatial coupling can be divided into quasi-rigid structural units that precisely correspond to natively folded domains. The high-PAE transition zones at the edges of dark green blocks in the heatmap correspond to flexible linker regions between domains, serving as an important complement to traditional sequence-alignment-based domain assignment methods, particularly offering higher precision for structurally rigid units with low sequence conservation.
3.3 Assessing the Authenticity of Protein Complex Interactions
In protein complex prediction, the PAE matrix is a key tool for eliminating false-positive interactions. In AlphaFold series outputs, if the average PAE value across the inter-chain interface region is below 10 Å, and the ipTM score and combined scores reach reliable thresholds, the predicted interface binding mode of the complex is considered reliable. If the inter-chain region exhibits large areas of high PAE values (>15 Å), this suggests that the interaction is most likely a false-positive result from random fitting and should not be used for interaction mechanism analysis.
3.4 Predicting Intrinsic Protein Dynamics and Flexibility
Studies have shown that the PAE matrix correlates significantly with the residue distance fluctuation matrix from molecular dynamics simulations and with B-factors. Low-PAE regions correspond to rigid, conserved regions, while high-PAE regions correspond to natively flexible regions. This enables the PAE matrix to be used for low-cost, rapid prediction of protein dynamic conformational features, suitable for protein flexibility analysis and conformational change mechanism studies.
3.5 Guiding Precision Protein Engineering
The PAE matrix can assist in screening effective engineering regions: low-PAE regions have high structural accuracy, and mutation designs based on these regions have strong reproducibility; high-PAE regions have high structural uncertainty and carry high engineering risks, and should be avoided or combined with experimental validation. Additionally, PAE can help identify stable domain scaffolds, providing a basis for de novo protein design.
IV. Core Considerations for Using the PAE Matrix
4.1 PAE Cannot Replace Experimental Validation
The PAE matrix is an estimated error value from algorithmic prediction, not an experimental measurement. A high PAE definitely indicates an unreliable structure, but a low PAE does not absolutely equate to the true structure—a low PAE only indicates that the conformation is the optimal structure fitted by the algorithm and may still deviate from the in vivo true conformation. All key conclusions based on PAE must ultimately be validated through wet-lab experiments such as X-ray crystallography, cryo-electron microscopy, or mutagenesis assays.
4.2 Multi-Metric Integrated Interpretation
A single PAE matrix cannot comprehensively evaluate structural quality and must be combined with multiple metrics: pLDDT assesses local folding quality, PAE assesses global spatial reliability—these two form the core evaluation system; pTM-score can serve as a supplementary reference for overall model quality, and ipTM-score specifically evaluates interface confidence in complex predictions. These metrics are complementary and synergistic, constituting the standard analytical workflow in the field.”
V. Conclusion
The PAE matrix is a core global confidence metric of deep learning-based structure prediction tools such as the AlphaFold series, overcoming the limitation of pLDDT in only evaluating local folding. Its core interpretive logic can be summarized in three points:
First, PAE values represent the spatial positional error after residue alignment and are unrelated to residue physical distances;
Second, the low-PAE blocks along the heatmap diagonal correspond to rigid domains, and low-PAE cross-region blocks represent reliable domain packing and protein interactions;
Third, high-PAE regions are structurally uncertain zones, and all functional conclusions should avoid these regions.
Currently, the PAE matrix has been deeply applied in core scenarios including multi-domain protein conformational assessment, domain boundary delineation, complex interaction validation, dynamic flexibility prediction, and precision protein engineering. As structure prediction algorithms continue to evolve, the interpretive framework and application scenarios of PAE are still expanding. For researchers in structural biology and protein engineering, accurate interpretation of the PAE matrix is an essential foundational skill for distinguishing true from false predicted structures and ensuring the reliability of research conclusions.