Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
A Critical Examination of "State-of-the-Art" Claims in Materials GNNs: Benchmark Artifacts and Unstated Limitations
“State-of-the-art” (SOTA) claims in materials graph neural networks are typically grounded in marginal improvements on standard benchmarks, yet such gains often reflect properties of the evaluation protocol rather than advances in generalization. This analysis demonstrates that widely used practices—particularly random train–test splits, inconsistent baselines, and selective reporting—systematically inflate reported performance. At the same time, critical constraints, including interpolation-dominated evaluation, computational cost, limited extrapolation, and dataset bias, remain largely unacknowledged. Drawing on the 2017–2023 literature, the study shows that these artifacts and omissions distort how progress is measured and interpreted. By reframing SOTA as a context-dependent rather than universal claim, it proposes a shift toward multi-split, multi-metric, and reproducible evaluation that more accurately reflects real-world materials discovery conditions.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2024 | Article: 26

Are We Overfitting to the Materials Project? A Critique of Dataset Bias in Property Prediction Benchmarks
The Materials Project (MP) has become the de facto benchmark for machine learning models in computational materials property prediction. Its vast repository of density-functional-theory (DFT) data has enabled rapid progress in graph neural networks, representation learning, and transfer-learning approaches. Yet this critique demonstrates that the field is systematically overfitting to MP’s structural and compositional artifacts. Far from being a neutral or representative sample of chemical space, MP is heavily skewed toward thermodynamically stable compounds lying near the convex hull, high-symmetry crystal systems, near-equiatomic stoichiometries, and a narrow set of well-studied elemental combinations. Models that achieve state-of-the-art mean absolute errors on MP test sets therefore learn dataset-specific regularities rather than transferable physical principles. Six interlocking forms of bias are identified: stability bias, symmetry bias, composition bias, prototype bias, element-frequency bias, and relaxation bias. Additional absence of disordered structures compounds the problem. The consequences are severe. Reported generalization performance is illusory when evaluated only on MP-held-out data; dramatic performance drops occur on independent databases, low-symmetry subsets, dilute alloys, and experimental formation energies. State-of-the-art claims based solely on MP benchmarks therefore mislead the community, waste experimental validation resources, and retard genuine advances in extrapolation and disorder handling. Drawing exclusively on peer-reviewed analyses published between 2017 and 2025, this work argues that the community’s MP-centric evaluation culture has created a benchmark artifact that masquerades as scientific progress. Debiased sampling, cross-dataset validation, subgroup reporting, and multi-database training are shown to be essential corrective measures. Until these practices become standard, claims of robust materials AI will remain premature. The critique concludes with concrete recommendations for benchmark designers, researchers, and journal editors to restore scientific integrity to property-prediction evaluation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2025 | Article: 47
Filters
Clear All





Access type