Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Uncertainty Quantification for ML Interatomic Potentials: A Review of Methods, Hidden Assumptions, and Unresolved Questions
Uncertainty quantification (UQ) has become indispensable for the trustworthy deployment of machine learning interatomic potentials (MLIPs) in materials science and molecular modeling, where predictions of energies, forces, and derived properties directly inform high-stakes decisions in materials discovery, long-time-scale molecular dynamics, and autonomous design workflows. Without reliable uncertainty estimates, MLIPs risk propagating errors that compromise simulation stability, mislead experimental prioritization, or produce unphysical results in extrapolation regimes critical to novel alloy or molecular discovery. This review synthesizes the literature on UQ methods specifically developed for or applied to MLIPs, drawing exclusively from the compiled reference set to provide a focused, critical overview of progress during this formative period. The scope is deliberately restricted to UQ techniques for interatomic potentials themselves (including GAP, DeepMD, ANI-series, SchNet-derived, and E(3)-equivariant models such as NequIP), excluding standalone ML property prediction unless the method directly supports force-field uncertainty. A systematic taxonomy organizes existing approaches into five methodological families—Bayesian and probabilistic methods, ensemble methods, Gaussian process and kernel methods, conformal prediction and frequentist methods, and heuristic and ad hoc methods—highlighting their distinct mathematical foundations and practical implementations in MLIP contexts. Hidden assumptions pervading these families are identified and dissected, including independence of atomic errors, Gaussianity of predictive distributions, homoscedasticity across chemical space, kernel-imposed smoothness in Gaussian processes, approximation quality in variational or Monte-Carlo inference, and exchangeability in conformal frameworks. These assumptions frequently remain unstated yet profoundly influence calibration and reliability when MLIPs are deployed in production simulations. Unresolved questions are articulated with precision: how to treat correlated uncertainties along molecular-dynamics trajectories, the absence of a true ground-truth uncertainty given DFT approximations, the prohibitive computational overhead of scalable UQ, evaluation under distribution shift, vectorial uncertainty for forces rather than scalar energies, detection of physical inconsistencies, and hierarchical fusion of model, data, and ab-initio uncertainties. Future outlook points toward integrated UQ-driven active learning, force-aware uncertainty representations, and hybrid methods that balance calibration, sharpness, and efficiency for next-generation autonomous materials engineering.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 July 2022 | Article: 9

Benchmarking Practices for ML Interatomic Potentials: A Critical Review of Methodological Pitfalls and What Was Missed (2017–2023)
Benchmarking has become central to the development of machine learning interatomic potentials (MLIPs), yet the epistemic reliability of reported comparisons remains insufficiently scrutinized. This review synthesizes prevailing practices and demonstrates that current evaluation protocols systematically misrepresent model progress. A coherent taxonomy of methodological failure emerges, spanning opaque data handling, structurally flawed train–test partitioning, restricted metric selection, weak baseline construction, limited reproducibility, and the near absence of extrapolation analysis. Under these conditions, widely cited performance indicators—such as sub-10 meV/atom energy MAE or sub-0.1 eV/Å force RMSE—primarily capture interpolation within constrained training distributions, offering limited insight into generalization, dynamical stability, or deployment viability. A related deficiency lies in the systematic exclusion of physically and computationally salient regimes, including long-range interactions, finite-temperature behavior, low-symmetry and disordered structures, calibrated uncertainty, defect-rich configurations, and explicit cost–accuracy trade-offs. Existing benchmark suites, including Materials Project–derived datasets, QM9 adaptations, COMP6, and bespoke collections, inherit these constraints, reinforcing an evaluative paradigm that privileges narrow optimization over robust, application-relevant performance. Recasting benchmark outcomes as contingent on methodological design rather than intrinsic model capability reveals how evaluation choices implicitly structure model rankings. In response, this work advances a set of directly implementable standards: diversified splitting strategies, distribution-aware multi-metric reporting, transparent baseline inclusion, controlled extrapolation regimes, complete reproducibility artifacts, and normalized cost accounting. Aligning benchmarking practice with these principles is necessary to transition from incremental leaderboard gains toward reliable and transferable interatomic potentials for materials discovery.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 January 2024 | Article: 28
Filters
Clear All





Access type Clear