Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
A Conceptual Theory of Measurement Validity for AI-Generated Materials Properties
The pervasive reliance on predictive accuracy metrics such as mean absolute error, root mean square error, and R² in materials artificial intelligence has created a fundamental misconception: that low prediction error equates to a valid measurement of a material’s property. This paper argues that accuracy alone is insufficient because an AI-generated property value may align closely with held-out test data yet fail to support the specific scientific or engineering inferences for which it is intended. Drawing on foundational measurement validity theory from psychometrics and the social sciences, the manuscript adapts these concepts to the unique context of AI-generated materials properties. It proposes a novel five-component conceptual theory of measurement validity tailored to machine-learning predictions of physical quantities such as band gaps, formation energies, and mechanical moduli. Five distinct dimensions of validity—construct, criterion, generalizability, robustness, and consequential—are articulated and illustrated with materials-specific scenarios. Finally, the framework offers concrete implications for authors, reviewers, and the broader materials informatics community, shifting validation practices from narrow accuracy reporting toward comprehensive evidence-based arguments that link predictions to intended uses. By distinguishing accuracy from validity, this conceptual framework aims to elevate the epistemological rigor of AI-driven materials discovery and design.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2024 | Article: 119

Toward Community-Agreed Metrics for Generative Materials Models: A Position on Inverse Design Reporting
Generative models for materials inverse design, including variational autoencoders, generative adversarial networks, and diffusion models, have rapidly emerged as powerful tools for proposing crystal structures conditioned on target properties. Yet, despite significant methodological progress, evaluation practices remain fragmented and inconsistent, limiting comparability and obscuring genuine scientific advancement. This review introduces a unified five-dimensional taxonomy—validity, novelty, diversity, property optimization, and stability—and uses it to diagnose systematic distortions in current reporting. Across the literature, metrics are defined heterogeneously, diversity is frequently omitted, stability is selectively assessed, and baseline comparisons are often absent. These inconsistencies inflate reported performance and prevent reliable benchmarking. In response, a minimal reporting standard is proposed, requiring transparent multi-metric evaluation, explicit dataset and split disclosure, baseline-controlled property assessment, distribution-level reporting, and computational-cost transparency. Establishing such community-agreed protocols is essential for transforming generative materials research into a reproducible, comparable, and cumulatively advancing field.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2024 | Article: 27

Defining "Novel Material" in Generative Models: A Boundary between Chemical Validity and Synthesizability
"Novel material" has become the central claim in nearly every paper introducing generative models for materials discovery. Yet the term is used with striking ambiguity. Authors routinely assert that their models have produced "novel materials" without clarifying whether this means the output lies outside the training distribution, satisfies basic chemical rules such as charge balance and realistic bond lengths, or meets the far stricter requirement of being synthesizable in a laboratory. This boundary/definitional article identifies three distinct meanings—distributional novelty, chemical validity, and synthesizability—and demonstrates how the current literature routinely conflates them. A generated crystal may be distributionally novel (absent from the training set) yet chemically invalid; it may be chemically valid yet lie far above the convex hull and therefore remain unsynthesizable. Such conflation leads to overclaims that inflate expectations and hinder reproducible progress in inverse design. This analysis maps the boundary conditions for each meaning and proposes a hierarchical operational definition with six explicit levels. The framework requires authors to report novelty percentages at every level rather than a single vague statistic. Distributional novelty marks the first filter, chemical validity the second, and synthesizability the third, with synthesizability itself subdivided into thermodynamic, dynamical, kinetic, and experimental realizability layers. The article further examines boundary cases, gray zones, and implications for model evaluation and benchmark design. By replacing ambiguous rhetoric with precise, multi-level reporting, the proposed definition establishes a shared language for generative models in materials engineering and prevents the overinterpretation of computational outputs as laboratory-ready discoveries. Adoption of this operational framework will sharpen claims, improve comparability across studies, and ultimately accelerate the translation of generative predictions into experimentally validated materials.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2025 | Article: 45
Filters
Clear All





Access type