- An RMSD plateau describes distance from a reference; it does not prove adequate sampling.
- Convergence belongs to an observable, tolerance and timescale—not permanently to a trajectory.
- Effective sample size, independent replicates, reversible transitions and model validation must agree.
1. What RMSD can—and cannot—establish
RMSD is a distance from a reference after a specified fit. Its value changes with the reference, atom selection, weighting, periodic-boundary treatment, and fitted region. A flat trace may represent a stable basin, but it may equally reflect kinetic trapping within one basin. Distinct conformations can collapse to similar RMSD values, while a flexible terminus can dominate a whole-protein signal and obscure a binding-site rearrangement. RMSD is useful for detecting gross drift and alignment problems; by itself, it says nothing about equilibrium populations or reversible exploration of configuration space.[1,4,5]
2. Convergence belongs to an observable
The first question should be “what quantity must be estimated, to what precision?” A mean domain distance may stabilize long before a rare opening rate, a low-population pocket occupancy, or a global free energy. Predefine the target—such as a contact occupancy with an interval, a state-population difference, or a slow timescale—and examine cumulative estimates, blockwise distributions, state counts, and transition counts. The 2024 study by Ormeño and General illustrates that different properties of the same protein can reach practical convergence on different timescales.[11]
Global and local diagnostics should be kept separate. Osato and colleagues showed in 2026 that slowly interconverting ligand and side-chain torsions can create disagreement between repeats while escaping a global structural summary.[12] Their workflow is deliberately a necessary, not sufficient, diagnostic: adequate torsion transitions do not rule out slow hydration, alternative poses, or large-scale protein motion. A useful dashboard therefore combines global shape, mechanism-linked coordinates, and data-driven slow modes.
3. Count independent information, not saved frames
Successive MD frames are correlated. For an observable with autocorrelation \rho(k), the statistical inefficiency may be estimated as g\approx1+2\sum_{k\ge1}\rho(k), giving N_{eff}\approx N_{prod}/g. Chodera's equilibration procedure chooses the production start t_0 that maximizes the number of effectively uncorrelated samples, making the bias–variance trade-off explicit.[2] These quantities must be computed for relevant observables: a rapidly decorrelating energy trace cannot certify a slow conformational coordinate.
Uncertainty estimates should retain correlation, using block averages, time-block bootstraps, or hierarchical resampling across independent runs, with block-length sensitivity reported. Splitting one trajectory in half is a drift check, not proof against an unseen slow process. When effective sample size is small, writing coordinates more frequently adds storage rather than information.[1,2]
4. Replicates reveal fragile conclusions
Independent runs should vary seeds or velocities and, when scientifically justified, start from multiple equilibrated configurations. Agreement should be tested on distributions, populations, transitions, and the final estimand—not on visual overlap alone. Knapp and colleagues demonstrated substantial single-trajectory irreproducibility in two systems and suggested five to ten replicas as a practical starting point under their tested conditions.[3] A 2023 Communications Biology editorial requires at least three independent simulations with statistical analysis.[7] Neither number is a universal theorem; the required ensemble depends on between-run variance and the desired precision.
Many short trajectories are not automatically superior to one long trajectory. Replication improves variance estimation and parallel discovery, but trajectories shorter than the controlling barrier-crossing time may all remain trapped. Kinetic questions additionally require connected, preferably reversible transitions. A sensible allocation uses early replicas to locate variance and bottlenecks, then extends selected runs or changes the sampling strategy based on missing transitions.
5. Diagnose slow coordinates before adding bias
Collective variables should distinguish hypothesized states and retain mechanistic meaning: domain opening, gate torsions, contacts, ligand pose, or pocket hydration. PCA, tICA/VAMP, and clustering can complement those coordinates by exposing unanticipated slow structure. Compare independent runs in a common representation, including coverage, populations, round trips, and the time evolution of estimates. A crisp two-dimensional free-energy surface is not persuasive if the basis changes with the time window or if replicas occupy different regions.[4–6,10]
Enhanced sampling does not remove the need for convergence analysis. Biased methods require credible CVs, reproducible walkers, and validated reweighting; a hidden slow mode orthogonal to the bias can still prevent equilibration. Replica-exchange analyses should examine overlap, exchanges, and round trips. Adaptive sampling can expand exploration by reseeding unbiased trajectory segments from informative states, but thermodynamic and kinetic estimates must remain compatible with the adaptive design.[6,9,10]
6. MSMs extend observations only after validation
MSMs can combine short trajectories that share states and estimate metastable populations, pathways, and slow kinetics. Their conclusions depend on features, dimensionality reduction, state discretization, lag time, connectivity, and observed transition counts. At minimum, report implied-timescale behavior, a Chapman–Kolmogorov test, sensitivity to state definitions and lag time, and resampling or Bayesian intervals.[4,5] Kozlowski and Grubmüller found limited sampling to be the dominant uncertainty in many MSMs of four small proteins; model-internal Bayesian intervals could underestimate broader uncertainty from construction choices.[8] An MSM cannot recover a state or pathway for which the data contain no evidence.
Common misconceptions and boundaries
A flat RMSD is not a converged ensemble; many frames are not many independent samples; three replicas are a baseline rather than a guarantee; absence of an event does not distinguish stability from trapping; a smooth free-energy surface does not validate its CVs; and an MSM timescale is not credible without lag-time, CK, and sampling checks. No diagnostic based on finite output can exclude every slower unseen process. Sampling quality also cannot repair force-field, protonation, parameterization, or boundary-condition errors. Biased and nonequilibrium trajectories require the appropriate correction before they are interpreted as equilibrium populations or natural kinetics.
Conclusion
RMSD belongs near the beginning of a quality-control workflow, not at its end. A credible MD claim is supported by a chain of evidence: a defined estimand, a justified production region, correlation-aware effective sample size, agreement and uncertainty across independent runs, repeated visits to relevant states, valid reweighting for enhanced sampling, and kinetic validation for MSMs. When those checks disagree, the honest interpretation is often that the current data cannot yet separate physical stability from metastability or kinetic trapping.
Diagnostic workflow
Define the estimand
observable, states, tolerance, and whether the goal is an equilibrium average, population, free energy, or rate.
Audit the trajectory
PBC treatment, imaging, alignment, reference and atom selection, equilibration and production ranges, and thermodynamic anomalies.
Use multiple scales
RMSD plus mechanistic local coordinates, hydration/contacts/torsions, slow modes, and clustered states.
Estimate independent information
observable-specific t_0, autocorrelation or statistical inefficiency, N_{eff}, and block-length sensitivity.
Compare independent runs
distributions, state populations, transitions, and uncertainty intervals; investigate rather than silently discard outliers.
Look for negative evidence
one-way transitions, low-count torsions, nonoverlapping replicas, late new states, and drifting estimates.
Match the remedy to the bottleneck
extend or replicate for variance; bias a known slow coordinate; use adaptive sampling for exploration; validate an MSM for kinetics.
Report limits
per-run and pooled estimates, transition counts, effective samples, reweighting and model tests, plus slow processes that remain unresolved.
Verified sources
- Grossfield A, Zuckerman DM. Chapter 2 Quantifying Uncertainty and Sampling Quality in Biomolecular Simulations. *Annual Reports in Computational Chemistry*. 2009;5:23–48.2009 · DOI 10.1016/S1574-1400(09
- Chodera JD. A Simple Method for Automated Equilibration Detection in Molecular Simulations. *Journal of Chemical Theory and Computation*. 2016;12(4):1799–1805.2016 · DOI 10.1021/acs.jctc.5b00784
- Knapp B, Ospina L, Deane CM. Avoiding False Positive Conclusions in Molecular Simulation: The Importance of Replicas. *Journal of Chemical Theory and Computation*. 2018;14(12):6127–6138.2018 · DOI 10.1021/acs.jctc.8b00391
- Husic BE, Pande VS. Markov State Models: From an Art to a Science. *Journal of the American Chemical Society*. 2018;140(7):2386–2396.2018 · DOI 10.1021/jacs.7b12191
- Prinz J-H, Wu H, Sarich M, et al. Markov models of molecular kinetics: Generation and validation. *The Journal of Chemical Physics*. 2011;134(17):174105.2011 · DOI 10.1063/1.3565032
- Hénin J, Lelièvre T, Shirts MR, Valsson O, Delemotte L. Enhanced Sampling Methods for Molecular Dynamics Simulations [Article v1.0]. *Living Journal of Computational Molecular Science*. 2022;4(1):1583.2022 · DOI 10.33011/livecoms.4.1.1583
- Reliability and reproducibility checklist for molecular dynamics simulations. *Communications Biology*. 2023;6:268.2023 · DOI 10.1038/s42003-023-04653-0
- Kozlowski N, Grubmüller H. Uncertainties in Markov State Models of Small Proteins. *Journal of Chemical Theory and Computation*. 2023;19(16):5516–5524.2023 · DOI 10.1021/acs.jctc.3c00372
- Kleiman DE, Nadeem H, Shukla D. Adaptive Sampling Methods for Molecular Dynamics in the Era of Machine Learning. *The Journal of Physical Chemistry B*. 2023;127(50):10669–10681.2023 · DOI 10.1021/acs.jpcb.3c04843
- Fu H, Bian H, Shao X, Cai W. Collective Variable-Based Enhanced Sampling: From Human Learning to Machine Learning. *The Journal of Physical Chemistry Letters*. 2024;15(6):1774–1783.2024 · DOI 10.1021/acs.jpclett.3c03542
- Ormeño F, General IJ. Convergence and equilibrium in molecular dynamics simulations. *Communications Chemistry*. 2024;7:26.2024 · DOI 10.1038/s42004-024-01114-5
- Osato M, Dabbous T, Mobley DL. An Automated Workflow for Diagnosing Sampling Issues Caused by Slow Torsional Motions in Molecular Simulations. *Journal of Chemical Information and Modeling*. 2026;66(10):5580–5594.2026 · DOI 10.1021/acs.jcim.6c00134
Search updated 2026-08-16. This is an evidence-led narrative methods review, not a registered systematic review or meta-analysis; citations prioritise primary papers, official documentation and standards.
