Key judgements
  • MM/GBSA, RBFE and ABFE answer different questions; they are not a universal accuracy ladder.
  • RBFE is best aligned with local congeneric changes; ABFE expands chemical reach and the sampling/state error budget.
  • Replicates, overlap, cycle closure and experimental anchors matter more than extra decimal places.

1. Separate the observable from the estimator

MM/GBSA is an endpoint approximation. Snapshots are usually drawn from a complex MD trajectory, then molecular-mechanics and continuum-solvation terms are evaluated for the complex, receptor, and ligand; entropy is frequently omitted or estimated separately. RBFE and ABFE name the desired observables: a ligand-to-ligand ΔΔG and a standard binding ΔG, respectively. FEP, TI, BAR, and MBAR name ways to connect and analyze thermodynamic states. An RBFE campaign can therefore be analyzed with TI or MBAR, while an industrial “FEP campaign” commonly means the broader RBFE process rather than exponential averaging alone.[3,6,7]

2. Start from the decision and the chemical neighborhood

MM/GBSA is often the pragmatic choice when hundreds or thousands of docked poses need a second-pass score, or when the aim is to compare poses within one modeled system. RBFE is a closer match to questions such as whether a local substituent change improves affinity inside a series with a shared binding mode. ABFE becomes relevant when ligands cannot be connected by sensible local transformations, no reference ligand is available, or an absolute scale is central to the question. Forcing unrelated molecules into one fragile RBFE network, or treating MM/GBSA values as transferable experimental ΔG measurements, expands the method beyond its evidence base.[2,3,5]

3. MM/GBSA is valuable because it is fast and inspectable

MM/GBSA can reuse conventional trajectories and naturally supports pose comparisons, conformational stratification, and residue-level decomposition. A single-trajectory protocol may reduce variance through cancellation, but assumes limited reorganization between bound and unbound states. Its error budget includes implicit solvent, dielectric choice, force field, conformational sampling, entropy, and discrete binding-site waters. Implementations and parameter choices can change the ranking. A 2022 benchmark across soluble and membrane proteins showed strong sensitivity to system-specific parameter choices; this is evidence for target-level calibration, not for universal default accuracy.[2,5,8]

4. RBFE has its clearest role in congeneric lead optimization

RBFE transforms ligand A into ligand B in solvent and in the binding site, then subtracts the two alchemical legs through a thermodynamic cycle. The shared environment and common substructure enable useful error cancellation. Success depends on a defensible bound structure, consistent protonation and tautomer states, credible atom mappings, and a connected perturbation network. Warning signs include net-charge changes, scaffold hops, ring changes, large substitutions, buried-water displacement, multiple poses, and slow protein rearrangements. Separated-topology approaches broaden the set of feasible transformations, but do not erase sampling or force-field limitations. Large retrospective and prospective studies support RBFE within its domain while also documenting substantial target dependence and outliers.[1,4,9–12]

5. ABFE buys chemical reach at the cost of a larger error budget

ABFE commonly uses a double-decoupling cycle in which ligand interactions are removed in solvent and in the binding site, together with positional/orientational restraints and a standard-state correction. It avoids the need to atom-map every ligand to a close reference, making it attractive for chemically diverse hits, fragments, or focused post-docking refinement. “Absolute,” however, does not mean automatically more accurate. Decoupling exposes problems involving multiple poses, pocket-water exchange, apo/holo protein reorganization, protonation coupling, restraints, and finite-size or charge corrections. Recent work demonstrates ABFE on pharmaceutically relevant systems, but the quality of the unbound protein ensemble and the completeness of sampling remain decisive.[6]

6. Inputs and experiments often set the accuracy ceiling

Missing loops, metals and cofactors, crystallographic waters, membrane context, ligand microstates, and incompatible assay conditions can dominate the error. Agreement among replicate trajectories measures precision inside sampled states; it cannot prove that the pose, protonation state, or slow conformation was correct. A defensible report combines independent repeats, estimator uncertainty, λ-state overlap, forward/reverse convergence, cycle closure, and structural inspection. It also uses historical compounds measured in the same assay system for retrospective calibration. Benchmark studies published from 2023 to 2026 show why aggregate accuracy cannot be transferred to a new target without qualification: experimental affinity itself has reproducibility limits, force-field errors are target dependent, and private discovery datasets can be substantially harder than curated public sets.[3,7,9–12]

7. Use a staged strategy, not a winner-takes-all method

A practical campaign starts with structural and chemical-state review, then uses docking or short MD to establish plausible poses. MM/GBSA can provide broad, inexpensive triage and flag unstable models. RBFE should be reserved for high-value compounds that form a defensible local transformation network. ABFE should answer selected questions where no useful reference exists, chemical space is discontinuous, or an absolute estimate is genuinely needed. Each layer needs negative controls, known-active anchors, and an explicit exit condition for calculations that are not trustworthy. Computational budget is usually better spent on independent repeats, difficult transformations, and alternative state hypotheses than on maximizing compound count.[3,4,10,12]

Method-selection matrix

  • Dimension: Primary question · MM/GBSA (endpoint): Fast rescoring, pose/conformer comparison, interaction interpretation · RBFE (often FEP/TI/BAR/MBAR): Within-target ΔΔG and ranking for related ligands · ABFE (absolute alchemy): Standard binding ΔG; focused comparison across diverse scaffolds
  • Dimension: Minimum inputs · MM/GBSA (endpoint): Plausible complexes; consistent topology/microstates; MD snapshots · RBFE (often FEP/TI/BAR/MBAR): High-quality bound structure; coherent series; consistent poses/microstates; mapping and network · ABFE (absolute alchemy): Credible pose; ligand/protein microstates; apo/holo strategy; restraints and standard state
  • Dimension: Experimental anchor · MM/GBSA (endpoint): Strongly recommended for parameter and threshold calibration · RBFE (often FEP/TI/BAR/MBAR): Strongly recommended as a retrospective test and network anchor · ABFE (absolute alchemy): Not required by theory, but protocol validation on known systems remains essential
  • Dimension: Relative cost · MM/GBSA (endpoint): Low–medium; inexpensive post-processing if MD already exists · RBFE (often FEP/TI/BAR/MBAR): High: two legs × λ states × network edges × repeats · ABFE (absolute alchemy): Very high: double decoupling, restraints/corrections, states, and repeats per ligand
  • Dimension: Defensible accuracy claim · MM/GBSA (endpoint): System dependent; useful for enrichment and trends, not guaranteed absolute ΔG · RBFE (often FEP/TI/BAR/MBAR): Often preferred for in-domain congeneric ranking; uncertainty remains edge specific · ABFE (absolute alchemy): Absolute in definition, but more exposed to slow sampling and systematic errors
  • Dimension: Essential diagnostics · MM/GBSA (endpoint): Snapshot convergence, block stability, dielectric/entropy sensitivity, pose consistency · RBFE (often FEP/TI/BAR/MBAR): Independent repeats, overlap, forward/reverse convergence, cycle closure, structural drift · ABFE (absolute alchemy): Independent repeats, coupling overlap, restraint stability, pose populations, apo/holo coverage
  • Dimension: Typical failure modes · MM/GBSA (endpoint): Missing entropy/waters; reorganization; tuning leakage; overinterpreted decomposition · RBFE (often FEP/TI/BAR/MBAR): Dissimilarity/charge changes; bad mapping; multiple poses; buried waters; slow reorganization · ABFE (absolute alchemy): Unknown pose; missing apo state; multiple microstates; standard-state/charge errors; ligand escape
  • Dimension: Best deliverable · MM/GBSA (endpoint): Ranking plus sensitivity analysis and interaction hypotheses · RBFE (often FEP/TI/BAR/MBAR): Networked ΔΔG with uncertainty, edge audit, and experimental comparison · ABFE (absolute alchemy): ΔG with state/restraint definitions, convergence evidence, and systematic-error discussion

Conclusion

The most reliable rule is not “use the most rigorous method the budget allows.” It is to match the observable, chemical similarity, structural confidence, and validation evidence. MM/GBSA serves broad and transparent triage; RBFE serves high-value ranking within a coherent series; ABFE addresses selected questions that local relative transformations cannot cover. All three require microstate review, independent repeats, and experimental anchors. Their value lies in reducing avoidable synthesis and inference errors—not in producing a single number with more decimal places.

DECISION GUIDE

Decision checklist

Check 1

Is the decision library enrichment, congeneric ranking, or a standard binding ΔG?

Check 2

Are same-assay historical activities and at least one credible complex structure available?

Check 3

Do ligands share a binding mode, charge state, and substantial common scaffold? If yes, test RBFE; if no, avoid forcing a single-topology map.

Check 4

Have protonation, tautomer, conformer, water, metal/cofactor, and membrane states been enumerated and recorded?

Check 5

Does the budget include independent repeats, difficult-edge extensions, and failed-run recovery?

Check 6

Are overlap, convergence, cycle-closure, and structural acceptance criteria defined before seeing results?

Check 7

Are experimental endpoints compatible, with IC50, Ki, and Kd distinguished rather than pooled uncritically?

Check 8

Will the report expose estimates, intervals, domain assumptions, and indeterminate cases together?

BOUNDARIES

Interpretive boundaries to retain

  • Per-residue MM/GBSA decomposition is attribution inside a model, not an experimentally separable residue binding free energy.
  • A small MBAR/FEP statistical error does not include a wrong pose, missing microstate, force-field bias, or experimental uncertainty.
  • Network-derived RBFE “DG” values are defined up to an additive offset; without an experimental anchor they are not measured absolute binding free energies.
  • Formal rigor in ABFE does not automatically solve apo/holo rearrangement, protonation coupling, or multiple binding poses.
  • Scores or ΔG values from different targets, assays, software, or force fields should not be merged into one universal leaderboard.
  • Benchmarks are affected by dataset selection, retrospective remediation, and publication bias; vendor-led and open benchmarks should be read together.
REFERENCES

Verified sources

  1. Wang L, Wu Y, Deng Y, et al. Accurate and Reliable Prediction of Relative Ligand Binding Potency in Prospective Drug Discovery by Way of a Modern Free-Energy Calculation Protocol and Force Field. *J Am Chem Soc.* 2015;137(7):2695–2703.2015 · DOI 10.1021/ja512751q
  2. Wang E, Sun H, Wang J, et al. End-Point Binding Free Energy Calculation with MM/PBSA and MM/GBSA: Strategies and Applications in Drug Design. *Chem Rev.* 2019;119(16):9478–9508.2019 · DOI 10.1021/acs.chemrev.9b00055
  3. Mey ASJS, Allen BK, Bruce Macdonald HE, et al. Best Practices for Alchemical Free Energy Calculations [Article v1.0]. *Living J Comput Mol Sci.* 2020;2(1).2020 · DOI 10.33011/livecoms.2.1.18378
  4. Schindler CEM, Baumann H, Blum A, et al. Large-Scale Assessment of Binding Free Energy Calculations in Active Drug Discovery Projects. *J Chem Inf Model.* 2020;60(11):5457–5474.2020 · DOI 10.1021/acs.jcim.0c00900
  5. Tuccinardi T. What is the current value of MM/PBSA and MM/GBSA methods in drug discovery? *Expert Opin Drug Discov.* 2021;16(11):1233–1237.2021 · DOI 10.1080/17460441.2021.1942836
  6. Khalak Y, Tresadern G, Aldeghi M, et al. Alchemical absolute protein–ligand binding free energies for drug design. *Chem Sci.* 2021;12(41):13958–13971.2021 · DOI 10.1039/D1SC03472C
  7. Wade AD, Bhati AP, Wan S, Coveney PV. Alchemical Free Energy Estimators and Molecular Dynamics Engines: Accuracy, Precision, and Reproducibility. *J Chem Theory Comput.* 2022;18(6):3972–3987.2022 · DOI 10.1021/acs.jctc.2c00114
  8. Wang S, Sun X, Cui W, Yuan S. MM/PB(GB)SA benchmarks on soluble proteins and membrane proteins. *Front Pharmacol.* 2022;13:1018351.2022 · DOI 10.3389/fphar.2022.1018351
  9. Ross GA, Lu C, Scarabelli G, et al. The maximal and current accuracy of rigorous protein-ligand binding free energy calculations. *Commun Chem.* 2023;6:222.2023 · DOI 10.1038/s42004-023-01019-9
  10. Hahn DF, Bayly CI, Boby ML, et al. Best Practices for Constructing, Preparing, and Evaluating Protein-Ligand Binding Affinity Benchmarks [Article v1.0]. *Living J Comput Mol Sci.* 2022;4(1):1497.2022 · DOI 10.33011/livecoms.4.1.1497
  11. Hahn DF, Gapsys V, de Groot BL, Mobley DL, Tresadern G. Current State of Open Source Force Fields in Protein–Ligand Binding Affinity Predictions. *J Chem Inf Model.* 2024;64(13):5063–5076.2024 · DOI 10.1021/acs.jcim.4c00417
  12. Baumann HM, Horton JT, Henry MM, et al. Large-Scale Collaborative Assessment of Binding Free Energy Calculations for Drug Discovery Using OpenFE. *J Chem Inf Model.* 2026;66(11):6429–6452.2026 · DOI 10.1021/acs.jcim.6c00089

Search updated 2026-08-16. This is an evidence-led narrative methods review, not a registered systematic review or meta-analysis; citations prioritise primary papers, official documentation and standards.