Key judgements
  • pLDDT primarily describes local confidence; PAE describes uncertainty in relative placement.
  • Predicted coordinates support structural hypotheses, not direct measurements of dynamics, affinity or mutation effects.
  • Pocket, interface and variant claims require biochemical context and task-matched external validation.

1. Prediction targets and biological claims are different objects

AlphaFold2 produces a static structural hypothesis together with internal error estimates. ESMFold replaces the multiple-sequence-alignment search with representations learned by a protein language model, enabling much faster single-sequence inference.[1,2] Their central benchmarks concern geometric agreement with experimental structures, not direct measurement of folding pathways, solution ensembles, free energies, or activity. AlphaFold3 broadens joint modelling to proteins, nucleic acids, selected ligands, ions, and modifications, yet its report documents stereochemical violations, hallucinations, missing dynamics, and lower accuracy for some target classes.[5] The first practical question is therefore not “Does the model look folded?” but “Which claim will these coordinates be asked to support?” Domain annotation, construct design, an interface hypothesis, docking, and variant interpretation each impose a different evidence threshold.

2. pLDDT is local confidence, not a whole-model certificate

pLDDT estimates the local lDDT-Cα accuracy around each residue. AlphaFold DB offers useful heuristics: values above 90 often indicate high local accuracy, 70–90 usually support a good backbone model, 50–70 warrant caution, and coordinates below 50 generally should not receive atomistic interpretation.[12] These bands are not guarantees. In an independent comparison with crystallographic maps, residues above 90 had a median Cα discrepancy of about 0.6 Å, but roughly 10% still differed by more than 2 Å; side-chain and environment-dependent details were less reliable.[4] Low pLDDT may reflect disorder, conditional folding, a missing partner, shallow evolutionary information, or model uncertainty. High pLDDT does not establish that the conformation dominates in vivo. Although downloaded files often store pLDDT in the B-factor field, it is not an experimental B factor and must not be interpreted as one.

3. Read PAE as a map of relationships

PAE(x,y) is the expected positional error at x when the model is aligned on y. Low diagonal blocks often identify internally coherent domains. If two domains each have high pLDDT but their cross-domain PAE is high, the defensible conclusion is that the domains may be individually useful while their relative orientation is unresolved. PAE is asymmetric, and collapsing it to one chain-average number destroys much of its purpose.[1,12] For complexes, examine inter-chain PAE, interface-residue pLDDT, ipTM/pTM, convergence across ranked models, interface chemistry, and plausible stoichiometry. AlphaFold2 produced many acceptable heterodimer models in an optimized-MSA benchmark, but incorrect interfaces remained.[11] A confident model of a supplied pair is not evidence that the proteins meet, coexist, or interact physiologically.

4. Accurate global folds do not guarantee docking-ready pockets

Ligand recognition depends on Å-scale side-chain rotamers, water, metals, protonation, tautomers, and induced fit. A small global RMSD and high pocket pLDDT can coexist with a docking-critical local error. In a PDBbind redocking benchmark, naïve AlphaFold2 targets achieved 17% pose success versus 41% for cognate crystal structures; removing low-confidence regions and allowing side-chain flexibility improved performance.[9] Yet prospective ultra-large-library screens against unrefined AlphaFold2 models of the σ2 and 5-HT2A receptors produced hit rates and affinity distributions comparable to screens against experimental structures, with experimental and cryo-EM follow-up supporting part of the campaign.[10] The lesson is target-specific, not universal: predicted pockets can be productive starting points when preparation, controls, and testing are strong. An AlphaFold3 protein–ligand pose is still not a Kd, Ki, ΔG, or cellular efficacy measurement.

5. One confident structure is not an ensemble

Functional proteins switch among states. Default AlphaFold2 usually returns one dominant-looking conformation, and AlphaFold DB explicitly states that predictions are not samples from a Boltzmann distribution and do not report relative state probabilities.[12] Reducing MSA depth allowed alternative, experimentally known states to be generated for several transporters and receptors, demonstrating a useful way to propose conformational candidates.[6] It did not turn model frequency into equilibrium population, and the authors required experimental validation of physiological relevance. Low pLDDT is therefore not RMSF, a state appearing in 70 of 100 predictions is not a 70% population, and failure to generate an open state is not evidence that opening cannot occur. Ensemble or kinetic claims require time-resolved or population-sensitive experiments, or simulations with appropriate sampling and convergence analysis.

6. Variant interpretation needs a stricter standard

Separately predicting a wild type and a point mutant, superposing their top-ranked models, and narrating a small displacement as disease mechanism often exceeds model resolution. AlphaFold DB marks mutation-effect prediction as unvalidated, and Buel and Walters showed examples in which destabilizing variants retained a confident, nearly unchanged predicted fold.[7,12] A large-scale study reached a more nuanced result: using an effective-strain statistic rather than visual comparison, it found aggregate agreement for localized deformation across 3,901 experimental/predicted structure pairs and proposed reliability filters.[8] This supports the possibility of task-specific, population-level signal; it does not validate a mechanistic structural change for every variant or provide ΔΔG, affinity loss, function, or clinical pathogenicity. Individual variants require conservation and population evidence, phenotype-aware models, and direct stability, binding, or functional measurements.

Decision checklist

  • The claim is restricted to regions whose local pLDDT and relevant PAE relationships support it.
  • Low pLDDT has not been equated automatically with disorder, and high pLDDT has not been equated with the physiological state.
  • Complexes are assessed with inter-chain PAE, interface metrics, model convergence, stoichiometry, and orthogonal interaction evidence.
  • Ligand work includes pocket preparation, protonation/tautomer choices, metals/waters, known-ligand controls, and experimental testing.
  • A wild-type–mutant coordinate difference is not reported as stability, affinity, function, or pathogenicity.
  • Dynamic language is supported by ensemble or time-resolved evidence, not confidence values or prediction frequency.
  • Conclusions say “supports a structural hypothesis” or “prioritizes a test,” not “proves the mechanism.”

Common misconceptions

  • Misconception: “Mean pLDDT is 92, so the entire model is ready to use.” · Better interpretation: The claim may depend on a weak loop, uncertain domain placement, or incorrect side chain; inspect residue confidence and PAE.
  • Misconception: “Low pLDDT is flexibility and can be treated as RMSF.” · Better interpretation: It is predicted structural error; disorder, missing context, and insufficient information can all lower it.
  • Misconception: “Both domains are blue, so their orientation is correct.” · Better interpretation: Domain pLDDT and inter-domain PAE answer different questions.
  • Misconception: “High ipTM proves a cellular interaction.” · Better interpretation: It scores a supplied complex model, not expression, colocalization, competition, or physiology.
  • Misconception: “Putting a ligand into AF3 yields affinity.” · Better interpretation: Joint structure prediction is not thermodynamic measurement.
  • Misconception: “A wild-type–mutant shift explains pathogenicity.” · Better interpretation: The shift may be below model error and requires task-specific and experimental evidence.

Conclusion

The strongest role of a predicted structure is to turn an unknown into a spatial, falsifiable hypothesis. pLDDT concerns local accuracy, PAE concerns relative placement, and interface scores concern the credibility of a proposed assembly. None is a proxy experiment for dynamics, affinity, or mutation effect. Claims become reliable when they are kept at the correct structural level, placed back into molecular context, checked across models, and validated with evidence matched to the task.

DECISION GUIDE

Confidence-reading workflow

Define the claim

Specify whether the endpoint is a domain boundary, local geometry, domain arrangement, interface, ligand pose, conformational state, or mutation effect.

Audit the input and predictor

Record sequence/isoform, species, truncations, oligomeric state, model version, templates/MSA, seeds, and included ligands, ions, or modifications.

Inspect residue-level pLDDT

Map confidence onto the exact residues that carry the interpretation; do not let a chain average hide a weak loop or pocket.

Read PAE by blocks

Separate within-domain, between-domain, and between-chain confidence and report local confidence independently from relative-placement uncertainty.

Test model consistency

Compare seeds, ranks, or model instances for topology, interface, pocket, and key side chains instead of selecting only the top score.

Restore biochemical context

Check cofactors, metals, membrane, modifications, pH, protonation, water, ligand-induced states, and experimental stoichiometry.

Validate the intended task

Use experimental structures, cross-links or mutations, binding measurements, enrichment benchmarks, or independent functional data at the same level as the claim.

BOUNDARIES

Interpretive boundaries to retain

  • Claim: Fold, domains, and conserved-residue localization · Defensible use: Often useful as a hypothesis · Minimum additional evidence: Local high pLDDT, low within-domain PAE, consistency with homologous or experimental data
  • Claim: Relative domain placement or protein interface · Defensible use: Conditional · Minimum additional evidence: Low cross-domain/chain PAE, interface metrics, model replication, orthogonal interaction evidence
  • Claim: Small-molecule pose or druggable pocket · Defensible use: Candidate hypothesis · Minimum additional evidence: Pocket refinement, known-ligand benchmark, multiple states, experimental structure or binding assay
  • Claim: State populations, rates, or allosteric pathways · Defensible use: Not supplied by one prediction · Minimum additional evidence: Ensemble-sensitive experiments or validated kinetic sampling
  • Claim: Mutation ΔΔG, affinity change, function, or pathogenicity · Defensible use: Not proven by coordinate differences · Minimum additional evidence: Stability/binding/function experiments and externally validated task models
REFERENCES

Verified sources

  1. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. *Nature*. 2021;596(7873):583–589.2021 · DOI 10.1038/s41586-021-03819-2
  2. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. *Science*. 2023;379(6637):1123–1130.2023 · DOI 10.1126/science.ade2574
  3. Akdel M, Pires DEV, Porta Pardo E, et al. A structural biology community assessment of AlphaFold2 applications. *Nature Structural & Molecular Biology*. 2022;29(11):1056–1067.2022 · DOI 10.1038/s41594-022-00849-w
  4. Terwilliger TC, Liebschner D, Croll TI, et al. AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. *Nature Methods*. 2024;21(1):110–116.2024 · DOI 10.1038/s41592-023-02087-4
  5. Abramson J, Adler J, Dunger J, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. *Nature*. 2024;630(8016):493–500.2024 · DOI 10.1038/s41586-024-07487-w
  6. del Alamo D, Sala D, Mchaourab HS, Meiler J. Sampling alternative conformational states of transporters and receptors with AlphaFold2. *eLife*. 2022;11:e75751.2022 · DOI 10.7554/eLife.75751
  7. Buel GR, Walters KJ. Can AlphaFold2 predict the impact of missense mutations on structure? *Nature Structural & Molecular Biology*. 2022;29(1):1–2.2022 · DOI 10.1038/s41594-021-00714-2
  8. McBride JM, Polev K, Abdirasulov A, et al. AlphaFold2 Can Predict Single-Mutation Effects. *Physical Review Letters*. 2023;131(21):218401.2023 · DOI 10.1103/PhysRevLett.131.218401
  9. Holcomb M, Chang Y-T, Goodsell DS, Forli S. Evaluation of AlphaFold2 structures as docking targets. *Protein Science*. 2023;32(1):e4530.2023 · DOI 10.1002/pro.4530
  10. Lyu J, Kapolka N, Gumpper R, et al. AlphaFold2 structures guide prospective ligand discovery. *Science*. 2024;384(6702):eadn6354.2024 · DOI 10.1126/science.adn6354
  11. Bryant P, Pozzati G, Elofsson A. Improved prediction of protein–protein interactions using AlphaFold2. *Nature Communications*. 2022;13:1265.2022 · DOI 10.1038/s41467-022-28865-w
  12. AlphaFold Protein Structure Database. FAQ: confidence, domain positions, complexes, unsupported use cases, and downloads. EMBL–EBI and Google DeepMind. Accessed 2026-08-16.2026

Search updated 2026-08-16. This is an evidence-led narrative methods review, not a registered systematic review or meta-analysis; citations prioritise primary papers, official documentation and standards.