Machine learning interatomic potentials (MLIPs) are increasingly used to accelerate atomistic simulation in materials science, yet their evaluation remains dominated by test-set error measured on held-out data drawn from the same distribution as the training set. Although this practice often yields low mean absolute errors for energies and forces, it provides limited evidence that a model will remain reliable when applied to the distribution shifts that define real deployment settings, including new compositions, defect structures, elevated temperatures, and non-equilibrium trajectories. This article develops a conceptual framework that distinguishes transferability from in-distribution accuracy and defines it as the ability of an MLIP to preserve predictive fidelity within application-relevant tolerances under explicitly characterized shifts in the joint distribution of structures and quantum-mechanical labels. The framework identifies four canonical forms of shift in computational materials science—compositional, structural, thermodynamic, and dynamical—and shows why current benchmarking practices systematically obscure them. To address this limitation, the study proposes a shift-aware evaluation protocol built around five components: shift characterization, sensitivity analysis, extrapolation-distance estimation, robustness criteria, and standardized reporting. Within this framework, Maximum Mean Discrepancy is adapted as a quantitative diagnostic for pre-deployment assessment of train–target divergence, while transferability is evaluated through five complementary dimensions: accuracy under shift, graceful degradation, uncertainty alignment, physical consistency, and compositional extrapolation. By replacing the tacit IID assumption with an explicit framework for reasoning about distribution shift, this work offers a common basis for evaluating, reporting, and comparing MLIP transferability across benchmarks and deployment scenarios, thereby supporting more trustworthy AI-driven materials discovery.
Open benchmark suites for machine learning interatomic potentials should include systematic failure cases as a core evaluation component rather than relying solely on leaderboard rankings based on average error metrics. Leaderboard-only benchmarks can obscure critical weaknesses, encourage overfitting to curated test sets, and provide limited guidance for real-world deployment. In contrast, failure-inclusive benchmarks reveal model limitations across challenging regimes such as compositional complexity, structural disorder, thermodynamic extremes, dynamical instability, numerical breakdown, and extrapolation beyond the training distribution. This position paper argues that making failure cases mandatory would improve transparency, support informed model selection, guide data acquisition, and strengthen reproducibility in materials machine learning. A practical typology of failure modes and a set of benchmark design principles are proposed to help shift evaluation from competitive ranking toward diagnostic understanding. Redesigning open benchmarks in this way is essential for building trustworthy and scientifically useful ML potentials.
Phonon spectra offer a uniquely demanding benchmark for graph neural network (GNN) force fields because they interrogate the second-derivative structure of the potential energy surface that governs lattice dynamics, thermal transport, and vibrational stability in materials. Yet direct comparison between GNN-derived and density functional theory (DFT) phonon spectra is frequently compromised by spectral leakage introduced through finite supercell truncation, displacement amplitude selection, q-point undersampling, Fourier interpolation, and post-processing broadening. These numerical effects can either conceal genuine deficiencies in the learned force field or generate apparent discrepancies that do not reflect model behavior. This article develops a hierarchical, leakage-aware validation framework that addresses this problem through progressive levels of scrutiny. The framework begins with baseline agreement in energies and forces, then advances to phonon density of states validation, q-resolved dispersion analysis, and finally a reproducible multi-metric assessment of spectral similarity. Progression through the hierarchy is conditional rather than automatic, such that higher-level claims are only made once lower-level numerical stability and model fidelity have been established. To separate methodological artifact from true representational error, the framework embeds explicit diagnostics based on supercell convergence, displacement sweeps, q-mesh refinement, interpolation cross-checks, and residual spectral analysis. It further introduces a standardized reporting protocol designed to make phonon-based validation transparent, comparable, and reproducible across studies of machine-learning interatomic potentials. By formalizing leakage control as an integral part of validation rather than an afterthought, the framework closes a critical gap between high-fidelity DFT phonon workflows and contemporary ML force-field development, enabling more credible assessment of GNN transferability in computational materials science.
Machine learning interatomic potentials (MLIPs) are widely validated against density functional theory (DFT), yet this practice risks reproducing DFT’s systematic biases rather than true experimental behavior. We present a concise four-level hierarchical framework for rigorous experimental validation focused on phonon dispersion and elastic constants. Level 1 evaluates static room-temperature elastic constants; Level 2 assesses zone-center phonon frequencies (Raman/IR); Level 3 examines full phonon dispersion curves from inelastic neutron/X-ray scattering; and Level 4 tests temperature-dependent shifts and softening. For each level we provide key observables, simple comparison metrics, uncertainty-aware criteria, and typical failure modes. A practical validation protocol and minimum reporting standards (Level 1 on ≥3 materials, Level 2 on ≥1) are proposed. This framework complements DFT benchmarks by anchoring MLIPs in experimental reality, improving reliability for thermal, mechanical, and vibrational applications in materials science.
The de-facto standard loss function employed in training machine learning interatomic potentials consists of a weighted combination of the mean squared error on total energies and the mean squared error on atomic forces. Researchers routinely adjust the relative weights assigned to energy and force terms in an attempt to achieve an optimal balance between these two objectives. This practice rests on the core assumption that energy errors and force errors maintain a simple linear relationship across all relevant atomic configurations. The present critique demonstrates that this assumption collapses when the loss function is applied to anharmonic systems where large-amplitude vibrations thermal disorder or phase-transition pathways dominate material behaviour. Four distinct failure modes are identified: harmonic bias force overfitting energy drift and extrapolation collapse. Each mode arises directly from the mismatch between the loss-function design and the non-quadratic nature of anharmonic potential-energy surfaces. Harmonic bias occurs because the weighted loss preferentially rewards models that reproduce quadratic energy landscapes typical of small-displacement training data even when those models are later deployed at elevated temperatures. Force overfitting emerges when the high-dimensional force term receives excessive weight causing the model to memorise training-set force patterns that do not generalise to unexplored configurational space. Energy drift follows because the balance achieved at zero-kelvin conditions no longer holds once thermal fluctuations populate regions of the energy surface far from the training distribution. Extrapolation collapse completes the picture because the loss function provides no explicit penalty for predictions outside the narrow domain of the training data rendering the model unusable for high-temperature properties. These failures have direct consequences for the prediction of thermal transport coefficients phonon lifetimes and finite-temperature stability in materials ranging from high-entropy alloys to solid electrolytes. The critique concludes by outlining detection principles and mitigation strategies that move beyond the energy–force balance paradigm advocating instead for anharmonic-aware loss designs that explicitly incorporate temperature-dependent information and higher-order derivatives. Adoption of such designs is essential if machine learning force fields are to deliver reliable predictions for the thermally activated processes that govern real-world materials performance.