Machine learning interatomic potentials (MLIPs) are increasingly used to accelerate atomistic simulation in materials science, yet their evaluation remains dominated by test-set error measured on held-out data drawn from the same distribution as the training set. Although this practice often yields low mean absolute errors for energies and forces, it provides limited evidence that a model will remain reliable when applied to the distribution shifts that define real deployment settings, including new compositions, defect structures, elevated temperatures, and non-equilibrium trajectories. This article develops a conceptual framework that distinguishes transferability from in-distribution accuracy and defines it as the ability of an MLIP to preserve predictive fidelity within application-relevant tolerances under explicitly characterized shifts in the joint distribution of structures and quantum-mechanical labels. The framework identifies four canonical forms of shift in computational materials science—compositional, structural, thermodynamic, and dynamical—and shows why current benchmarking practices systematically obscure them. To address this limitation, the study proposes a shift-aware evaluation protocol built around five components: shift characterization, sensitivity analysis, extrapolation-distance estimation, robustness criteria, and standardized reporting. Within this framework, Maximum Mean Discrepancy is adapted as a quantitative diagnostic for pre-deployment assessment of train–target divergence, while transferability is evaluated through five complementary dimensions: accuracy under shift, graceful degradation, uncertainty alignment, physical consistency, and compositional extrapolation. By replacing the tacit IID assumption with an explicit framework for reasoning about distribution shift, this work offers a common basis for evaluating, reporting, and comparing MLIP transferability across benchmarks and deployment scenarios, thereby supporting more trustworthy AI-driven materials discovery.
Open benchmark suites for machine learning interatomic potentials should include systematic failure cases as a core evaluation component rather than relying solely on leaderboard rankings based on average error metrics. Leaderboard-only benchmarks can obscure critical weaknesses, encourage overfitting to curated test sets, and provide limited guidance for real-world deployment. In contrast, failure-inclusive benchmarks reveal model limitations across challenging regimes such as compositional complexity, structural disorder, thermodynamic extremes, dynamical instability, numerical breakdown, and extrapolation beyond the training distribution. This position paper argues that making failure cases mandatory would improve transparency, support informed model selection, guide data acquisition, and strengthen reproducibility in materials machine learning. A practical typology of failure modes and a set of benchmark design principles are proposed to help shift evaluation from competitive ranking toward diagnostic understanding. Redesigning open benchmarks in this way is essential for building trustworthy and scientifically useful ML potentials.
Uncertainty quantification (UQ) has become indispensable for the trustworthy deployment of machine learning interatomic potentials (MLIPs) in materials science and molecular modeling, where predictions of energies, forces, and derived properties directly inform high-stakes decisions in materials discovery, long-time-scale molecular dynamics, and autonomous design workflows. Without reliable uncertainty estimates, MLIPs risk propagating errors that compromise simulation stability, mislead experimental prioritization, or produce unphysical results in extrapolation regimes critical to novel alloy or molecular discovery. This review synthesizes the literature on UQ methods specifically developed for or applied to MLIPs, drawing exclusively from the compiled reference set to provide a focused, critical overview of progress during this formative period. The scope is deliberately restricted to UQ techniques for interatomic potentials themselves (including GAP, DeepMD, ANI-series, SchNet-derived, and E(3)-equivariant models such as NequIP), excluding standalone ML property prediction unless the method directly supports force-field uncertainty. A systematic taxonomy organizes existing approaches into five methodological families—Bayesian and probabilistic methods, ensemble methods, Gaussian process and kernel methods, conformal prediction and frequentist methods, and heuristic and ad hoc methods—highlighting their distinct mathematical foundations and practical implementations in MLIP contexts. Hidden assumptions pervading these families are identified and dissected, including independence of atomic errors, Gaussianity of predictive distributions, homoscedasticity across chemical space, kernel-imposed smoothness in Gaussian processes, approximation quality in variational or Monte-Carlo inference, and exchangeability in conformal frameworks. These assumptions frequently remain unstated yet profoundly influence calibration and reliability when MLIPs are deployed in production simulations. Unresolved questions are articulated with precision: how to treat correlated uncertainties along molecular-dynamics trajectories, the absence of a true ground-truth uncertainty given DFT approximations, the prohibitive computational overhead of scalable UQ, evaluation under distribution shift, vectorial uncertainty for forces rather than scalar energies, detection of physical inconsistencies, and hierarchical fusion of model, data, and ab-initio uncertainties. Future outlook points toward integrated UQ-driven active learning, force-aware uncertainty representations, and hybrid methods that balance calibration, sharpness, and efficiency for next-generation autonomous materials engineering.
Phonon spectra offer a uniquely demanding benchmark for graph neural network (GNN) force fields because they interrogate the second-derivative structure of the potential energy surface that governs lattice dynamics, thermal transport, and vibrational stability in materials. Yet direct comparison between GNN-derived and density functional theory (DFT) phonon spectra is frequently compromised by spectral leakage introduced through finite supercell truncation, displacement amplitude selection, q-point undersampling, Fourier interpolation, and post-processing broadening. These numerical effects can either conceal genuine deficiencies in the learned force field or generate apparent discrepancies that do not reflect model behavior. This article develops a hierarchical, leakage-aware validation framework that addresses this problem through progressive levels of scrutiny. The framework begins with baseline agreement in energies and forces, then advances to phonon density of states validation, q-resolved dispersion analysis, and finally a reproducible multi-metric assessment of spectral similarity. Progression through the hierarchy is conditional rather than automatic, such that higher-level claims are only made once lower-level numerical stability and model fidelity have been established. To separate methodological artifact from true representational error, the framework embeds explicit diagnostics based on supercell convergence, displacement sweeps, q-mesh refinement, interpolation cross-checks, and residual spectral analysis. It further introduces a standardized reporting protocol designed to make phonon-based validation transparent, comparable, and reproducible across studies of machine-learning interatomic potentials. By formalizing leakage control as an integral part of validation rather than an afterthought, the framework closes a critical gap between high-fidelity DFT phonon workflows and contemporary ML force-field development, enabling more credible assessment of GNN transferability in computational materials science.
Benchmarking has become central to the development of machine learning interatomic potentials (MLIPs), yet the epistemic reliability of reported comparisons remains insufficiently scrutinized. This review synthesizes prevailing practices and demonstrates that current evaluation protocols systematically misrepresent model progress. A coherent taxonomy of methodological failure emerges, spanning opaque data handling, structurally flawed train–test partitioning, restricted metric selection, weak baseline construction, limited reproducibility, and the near absence of extrapolation analysis. Under these conditions, widely cited performance indicators—such as sub-10 meV/atom energy MAE or sub-0.1 eV/Å force RMSE—primarily capture interpolation within constrained training distributions, offering limited insight into generalization, dynamical stability, or deployment viability. A related deficiency lies in the systematic exclusion of physically and computationally salient regimes, including long-range interactions, finite-temperature behavior, low-symmetry and disordered structures, calibrated uncertainty, defect-rich configurations, and explicit cost–accuracy trade-offs. Existing benchmark suites, including Materials Project–derived datasets, QM9 adaptations, COMP6, and bespoke collections, inherit these constraints, reinforcing an evaluative paradigm that privileges narrow optimization over robust, application-relevant performance. Recasting benchmark outcomes as contingent on methodological design rather than intrinsic model capability reveals how evaluation choices implicitly structure model rankings. In response, this work advances a set of directly implementable standards: diversified splitting strategies, distribution-aware multi-metric reporting, transparent baseline inclusion, controlled extrapolation regimes, complete reproducibility artifacts, and normalized cost accounting. Aligning benchmarking practice with these principles is necessary to transition from incremental leaderboard gains toward reliable and transferable interatomic potentials for materials discovery.
Machine learning interatomic potentials (MLIPs) are widely validated against density functional theory (DFT), yet this practice risks reproducing DFT’s systematic biases rather than true experimental behavior. We present a concise four-level hierarchical framework for rigorous experimental validation focused on phonon dispersion and elastic constants. Level 1 evaluates static room-temperature elastic constants; Level 2 assesses zone-center phonon frequencies (Raman/IR); Level 3 examines full phonon dispersion curves from inelastic neutron/X-ray scattering; and Level 4 tests temperature-dependent shifts and softening. For each level we provide key observables, simple comparison metrics, uncertainty-aware criteria, and typical failure modes. A practical validation protocol and minimum reporting standards (Level 1 on ≥3 materials, Level 2 on ≥1) are proposed. This framework complements DFT benchmarks by anchoring MLIPs in experimental reality, improving reliability for thermal, mechanical, and vibrational applications in materials science.
The de-facto standard loss function employed in training machine learning interatomic potentials consists of a weighted combination of the mean squared error on total energies and the mean squared error on atomic forces. Researchers routinely adjust the relative weights assigned to energy and force terms in an attempt to achieve an optimal balance between these two objectives. This practice rests on the core assumption that energy errors and force errors maintain a simple linear relationship across all relevant atomic configurations. The present critique demonstrates that this assumption collapses when the loss function is applied to anharmonic systems where large-amplitude vibrations thermal disorder or phase-transition pathways dominate material behaviour. Four distinct failure modes are identified: harmonic bias force overfitting energy drift and extrapolation collapse. Each mode arises directly from the mismatch between the loss-function design and the non-quadratic nature of anharmonic potential-energy surfaces. Harmonic bias occurs because the weighted loss preferentially rewards models that reproduce quadratic energy landscapes typical of small-displacement training data even when those models are later deployed at elevated temperatures. Force overfitting emerges when the high-dimensional force term receives excessive weight causing the model to memorise training-set force patterns that do not generalise to unexplored configurational space. Energy drift follows because the balance achieved at zero-kelvin conditions no longer holds once thermal fluctuations populate regions of the energy surface far from the training distribution. Extrapolation collapse completes the picture because the loss function provides no explicit penalty for predictions outside the narrow domain of the training data rendering the model unusable for high-temperature properties. These failures have direct consequences for the prediction of thermal transport coefficients phonon lifetimes and finite-temperature stability in materials ranging from high-entropy alloys to solid electrolytes. The critique concludes by outlining detection principles and mitigation strategies that move beyond the energy–force balance paradigm advocating instead for anharmonic-aware loss designs that explicitly incorporate temperature-dependent information and higher-order derivatives. Adoption of such designs is essential if machine learning force fields are to deliver reliable predictions for the thermally activated processes that govern real-world materials performance.