Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Benchmarking Practices in Materials Artificial Intelligence — What Is Measured and What Is Missed
Materials artificial intelligence (MAI) has revolutionized the discovery, design, and optimization of new materials by leveraging machine learning algorithms to analyze complex datasets and predict properties with high accuracy. However, the rapid proliferation of MAI tools has raised critical questions about benchmarking practices, which are essential for evaluating model performance, ensuring reproducibility, and addressing ethical concerns. This narrative review examines current benchmarking frameworks in MAI, highlighting what is effectively measured—such as predictive accuracy and computational efficiency—and what is often overlooked —such as data bias, interpretability, fairness, and ethical implications. Drawing on recent advances in frameworks such as JARVIS-Leaderboard and Matbench, the review discusses challenges in data quality, reproducibility, and the integration of explainable AI (XAI) methods. It also explores active learning strategies for optimizing materials discovery under limited data conditions and proposes directions for more inclusive and transparent benchmarking. By synthesizing insights from diverse studies, this review aims to guide future MAI research toward robust, equitable, and ethically sound practices that accelerate innovation while mitigating risks.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2026 | Article: 92

Why Most ML Potential Benchmarks Overestimate Real-World Accuracy on Disordered Alloys
Machine learning interatomic potentials have become central tools in computational materials engineering, promising accurate and scalable predictions of alloy properties. Benchmarks in the field routinely report impressive accuracies, such as energy mean absolute errors below 20 meV/atom and low force root-mean-square errors, leading many papers to conclude that their models generalize well to alloy systems. Yet when these same potentials are deployed on real-world disordered alloys—such as binary solid solutions or multi-principal-element high-entropy alloys—the accuracy collapses, often by an order of magnitude. The root cause lies in the benchmarking pipeline itself: training and testing occur almost exclusively on ordered crystal structures drawn from databases of intermetallic compounds, while disordered configurations, which dominate practical applications, are never evaluated. This critical critique identifies four primary sources of systematic overestimation. First, the near-universal reliance on ordered training and test sets creates an illusion of generalization that fails the moment configurational disorder is introduced. Second, random or composition-based splits on ordered data allow models to memorize recurring structural motifs rather than learn transferable physics. Third, even when special quasirandom structures are employed as proxies for disorder, their periodic nature and limited sampling of local environments produce a misleading “SQS mirage.” Fourth, energy-only metrics mask severe force-prediction failures that render molecular-dynamics simulations unstable in disordered systems. These benchmarking flaws have concrete consequences: overconfidence among practitioners, misallocation of experimental resources, and slowed progress in high-entropy alloy design. By drawing on the literature, this article demonstrates that the ordered-to-disordered gap is not an edge case but a fundamental and under-reported limitation. It proposes concrete reforms—mandatory disordered test sets, configurational sampling protocols, and separate reporting of ordered versus disordered performance—to restore credibility to ML potential benchmarks. Until these changes are adopted, claims of “alloy-ready” potentials should be treated with skepticism. The field must move beyond ordered-crystal comfort zones if machine-learned potentials are to deliver on their promise for the compositionally complex materials that define modern materials science.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2023 | Article: 19

Benchmarking Practices for ML Interatomic Potentials: A Critical Review of Methodological Pitfalls and What Was Missed (2017–2023)
Benchmarking has become central to the development of machine learning interatomic potentials (MLIPs), yet the epistemic reliability of reported comparisons remains insufficiently scrutinized. This review synthesizes prevailing practices and demonstrates that current evaluation protocols systematically misrepresent model progress. A coherent taxonomy of methodological failure emerges, spanning opaque data handling, structurally flawed train–test partitioning, restricted metric selection, weak baseline construction, limited reproducibility, and the near absence of extrapolation analysis. Under these conditions, widely cited performance indicators—such as sub-10 meV/atom energy MAE or sub-0.1 eV/Å force RMSE—primarily capture interpolation within constrained training distributions, offering limited insight into generalization, dynamical stability, or deployment viability. A related deficiency lies in the systematic exclusion of physically and computationally salient regimes, including long-range interactions, finite-temperature behavior, low-symmetry and disordered structures, calibrated uncertainty, defect-rich configurations, and explicit cost–accuracy trade-offs. Existing benchmark suites, including Materials Project–derived datasets, QM9 adaptations, COMP6, and bespoke collections, inherit these constraints, reinforcing an evaluative paradigm that privileges narrow optimization over robust, application-relevant performance. Recasting benchmark outcomes as contingent on methodological design rather than intrinsic model capability reveals how evaluation choices implicitly structure model rankings. In response, this work advances a set of directly implementable standards: diversified splitting strategies, distribution-aware multi-metric reporting, transparent baseline inclusion, controlled extrapolation regimes, complete reproducibility artifacts, and normalized cost accounting. Aligning benchmarking practice with these principles is necessary to transition from incremental leaderboard gains toward reliable and transferable interatomic potentials for materials discovery.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 January 2024 | Article: 28
Filters
Clear All





Access type