Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
The Illusion of Generalizability in Foundation Models for Materials Science: A Critique of Zero-Shot Property Prediction
Foundation models promise a transformative advance in materials science by enabling zero-shot property prediction across diverse chemistries, crystal structures, and physical properties, ostensibly eliminating the need for property-specific labelled datasets and expensive first-principles calculations. Proponents claim that models pre-trained on millions of structures can directly predict formation energies, band gaps, elastic moduli, thermal conductivities, and other attributes for entirely unseen materials without task-specific fine-tuning. This critique demonstrates that such claims largely reflect an illusion of generalizability. Reported zero-shot performance is systematically inflated by six pervasive artefacts: (1) test-set leakage of crystal structures or near-identical analogues from the pre-training corpus, (2) exploitation of strong inter-property correlations, (3) evaluation within the interpolation regime of the pre-training distribution, (4) benchmark bias inherent in widely reused datasets such as the Materials Project, (5) inconsistent pre-processing, data splits, and reporting practices, and (6) the absence of rigorous elementary baselines. Analysis of recent foundation-model literature (2023–2026) shows that zero-shot accuracies frequently collapse once these confounds are controlled and often fail to surpass simple statistical predictors such as k-nearest-neighbour regression on embeddings or linear models based on property correlations. The illusion is particularly acute in materials science due to the repeated reuse of identical structures across property databases and the dense network of physical correlations among computed quantities. Without stringent controls, overstated zero-shot claims risk misdirecting research resources, inflating expectations, and undermining trust in AI-driven materials discovery. This paper provides a diagnostic framework for identifying these artefacts, marshals supporting evidence from the contemporary literature, and proposes a six-point evaluation standard for credible zero-shot claims. Adoption of these practices is essential if foundation models are to deliver genuine advances in out-of-distribution generalization rather than repackaged statistical regularities.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2026 | Article: 63
Filters
Clear All





Access type