The integration of artificial intelligence (AI) into materials science has substantially accelerated property prediction and materials screening. Yet, the predominance of data-driven correlations has exposed a persistent epistemic gap between predictive success and the derivation of interpretable, generalizable design rules. This conceptual manuscript develops a theoretical framework for knowledge extraction in materials AI that explicitly addresses this gap by reframing the transition from correlations to design rules as a staged epistemic process rather than a by-product of model performance. Drawing on literature in materials informatics, data bias, and philosophy of science, the framework organizes knowledge extraction into four interconnected stages—Correlation Mapping, Bias Interrogation, Value Integration, and Rule Synthesis—linked through continuous epistemic validation. The model foregrounds epistemic agency, requiring explicit scrutiny of assumptions, biases, and value commitments before causal inference. Six propositions articulate the conditions under which AI-derived correlations may legitimately support prescriptive design claims, emphasizing reflexive feedback and epistemic governance. By conceptualizing knowledge extraction as a norm-governed process of justification, this work provides a theoretical scaffold for transforming AI outputs into scientifically defensible design rules, contributing to a more reliable and responsible epistemology of materials discovery.
The integration of artificial intelligence (AI) into materials science has ushered in an era of semi-autonomous systems that accelerate discovery through predictive modeling, high-throughput screening, and adaptive experimentation. These systems offer substantial promise for addressing global challenges in energy, sustainability, and advanced manufacturing; however, their reliance on data-driven inference introduces risks related to bias propagation, epistemic uncertainty, and misalignment with scientific values. Conventional approaches treat human oversight primarily as an external corrective mechanism—post hoc monitoring or intervention in response to model outputs. This paper proposes a conceptual reframing wherein human oversight is repositioned as an intrinsic element of system design. Rather than viewing control as supervision layered atop an autonomous core, oversight is conceptualized as deliberate architectural choices that embed human judgment into the foundational structure of semi-autonomous materials AI. Drawing on literature from materials informatics, data bias mitigation, explainable AI, and human-AI collaboration, the proposed framework delineates three interdependent dimensions: epistemic boundary-setting, value-aligned modulation, and adaptive reflexivity. This reframing shifts the discourse from mitigating human absence to engineering human presence, fostering systems that are inherently more robust, interpretable, and aligned with the normative goals of scientific inquiry. By reconceptualizing oversight as design, the framework offers a pathway to responsible integration of AI in materials discovery without presupposing full autonomy or diminishing human agency.
In the rapidly evolving field of materials artificial intelligence (AI), the prevailing emphasis on scaling data volumes and computational resources has driven significant advancements in predictive modeling and discovery processes. However, this conceptual manuscript interrogates the implicit assumption that larger scales invariably yield superior outcomes, positing instead that unchecked expansion introduces intricate interaction dynamics that undermine the integrity of materials informatics. Through an integrative analysis, we explore how escalating data scales interact with inherent biases, leading to amplified distortions in representational fidelity and epistemic reliability. The framework delineates trade-offs wherein quantitative abundance may erode qualitative depth, fostering feedback structures that perpetuate homogeneity in material explorations at the expense of diversity. Ethical reasoning underscores the epistemic implications, revealing how scale-driven approaches can inadvertently prioritize dominant paradigms, marginalizing underrepresented material classes and contexts. Systems-level insights highlight steering logics that balance scale with interpretive nuance, advocating for calibrated integrations that preserve domain-specific insights. This argument reframes scale not as an unequivocal virtue but as a contingent factor within broader conceptual interpretations, urging a reevaluation of priorities in applied AI for materials science to foster sustainable and equitable progress.
Materials exploration faces persistent challenges stemming from vast chemical spaces, high experimental costs, and inherent uncertainties in predictive models. While machine learning has accelerated property prediction and guided candidate selection, conventional approaches often treat uncertainty as a uniform metric within fixed acquisition strategies. This conceptual paper introduces uncertainty-conditioned experiment planning (UCEP) as a novel theoretical framework for AI-guided materials discovery. UCEP reframes experiment planning as a dynamic process conditioned on the multidimensional character of uncertainty, integrating epistemic and aleatoric components, data-related biases, and model limitations into the steering logic. Rather than relying on static acquisition functions, the framework emphasizes adaptive interaction dynamics between uncertainty characterization and planning decisions, enabling context-sensitive trade-offs between exploration, exploitation, and bias mitigation. Drawing on interpretive insights from materials informatics and uncertainty quantification literature, UCEP highlights systems-level feedback structures that can enhance epistemic robustness and scientific efficiency without presupposing empirical outcomes. The framework offers analytical implications for rethinking how AI systems interpret and respond to uncertainty in iterative discovery cycles, contributing to more reflective and integrative AI-assisted materials research.
The integration of artificial intelligence into materials science has accelerated property prediction, inverse design, and discovery pipelines. Yet, the reliability of resulting scientific claims remains vulnerable to distribution shifts—systematic differences between training and inference data distributions arising from variations in synthesis protocols, characterization instruments, environmental conditions, or sampling biases. This purely conceptual manuscript develops a novel theoretical framework for robust materials AI inference in the presence of such shifts. We posit that distribution shifts do not merely degrade predictive accuracy but fundamentally alter the epistemic status of scientific claims by introducing unaccounted covariances between material descriptors and latent generative processes. The framework reconceptualizes inference as a multi-layered epistemic process: (i) shift ontology delineation, (ii) value-laden alignment of data representations with domain invariants, and (iii) claim robustness via counterfactual stabilization. By synthesizing insights from materials informatics, machine learning theory on distribution shifts, and philosophical analyses of epistemic values in science, we argue that robust inference requires explicit modeling of shift-induced epistemic uncertainty rather than mitigation as a post hoc engineering concern. This theory provides a conceptual scaffold for evaluating the validity of AI-derived materials claims across heterogeneous datasets, advancing a shift from performance-centric to epistemically grounded AI deployment in materials science.
The integration of artificial intelligence (AI) and machine learning (ML) into materials science, often referred to as materials informatics or materials AI, has accelerated the discovery, design, and optimization of advanced materials. However, materials science frequently operates in small-data and sparse-regime conditions, where datasets are limited in size (often tens to hundreds of samples), high-dimensional, imbalanced, or sparsely populated due to the high cost, time, and complexity of experimental measurements and high-fidelity simulations. This narrative review synthesizes recent advances in methods tailored to these constraints, categorizing approaches at the data-source level (e.g., literature extraction, database construction, high-throughput workflows), algorithmic level (e.g., support vector machines, Gaussian process regression, ensemble models, imbalanced learning techniques), and strategic level (e.g., active learning, transfer learning). Key assumptions underlying these methods are examined, including similarity between source and target domains for transfer learning, representativeness of initial samples and reliable uncertainty quantification in active learning, and the validity of physical priors or inductive biases in physics-informed approaches. The review also addresses inherent limits, such as risks of overfitting, poor generalization beyond the training distribution, sensitivity to data quality and noise, challenges in uncertainty calibration, and dependence on domain expertise. By highlighting successful applications in property prediction, alloy design, and perovskite optimization, this work elucidates the current capabilities and boundaries of small-data and sparse-regime learning in materials AI, guiding researchers navigating data-limited environments.
The integration of physical principles into machine learning (ML) frameworks has emerged as a transformative approach in materials science, addressing the limitations of purely data-driven models by incorporating domain knowledge to enhance predictive accuracy, generalizability, and interpretability. This narrative review explores the conceptual taxonomies of physics-integrated ML methods, their applications in materials discovery and design, and the associated challenges in data bias and ethical considerations. Drawing on recent peer-reviewed literature, we classify physics-integration strategies such as physics-informed neural networks (PINNs), hybrid models combining ML with physical simulations, and constraint-based learning, and highlight their roles in solving complex problems such as material property prediction, microstructure analysis, and phase stability. We also examine how data biases in training datasets can propagate errors and inequities in model outputs, and discuss the ethical values underpinning the use of AI in scientific research, including transparency, accountability, and societal impact. The review underscores the potential of these methods to accelerate innovation in materials science while emphasizing the need for rigorous validation and interdisciplinary collaboration. By synthesizing current advancements, this article aims to provide a foundational understanding for researchers and practitioners, paving the way for future developments in this interdisciplinary field.
Materials artificial intelligence (MAI) has revolutionized the discovery, design, and optimization of new materials by leveraging machine learning algorithms to analyze complex datasets and predict properties with high accuracy. However, the rapid proliferation of MAI tools has raised critical questions about benchmarking practices, which are essential for evaluating model performance, ensuring reproducibility, and addressing ethical concerns. This narrative review examines current benchmarking frameworks in MAI, highlighting what is effectively measured—such as predictive accuracy and computational efficiency—and what is often overlooked —such as data bias, interpretability, fairness, and ethical implications. Drawing on recent advances in frameworks such as JARVIS-Leaderboard and Matbench, the review discusses challenges in data quality, reproducibility, and the integration of explainable AI (XAI) methods. It also explores active learning strategies for optimizing materials discovery under limited data conditions and proposes directions for more inclusive and transparent benchmarking. By synthesizing insights from diverse studies, this review aims to guide future MAI research toward robust, equitable, and ethically sound practices that accelerate innovation while mitigating risks.
Autonomous and semi-autonomous laboratories represent a transformative paradigm in materials science, integrating artificial intelligence, robotics, and high-throughput experimentation to accelerate discovery and optimization processes. This review examines the conceptual foundations of these systems, including closed-loop optimization, machine learning algorithms, and modular hardware architectures. We explore their applications in areas such as alloy development, perovskite synthesis, and nanoparticle engineering, highlighting successes that have reduced discovery timelines from years to days. However, we also critically assess associated risks, including data quality issues, algorithmic biases, ethical concerns in resource allocation, and potential safety hazards from unsupervised operations. Drawing on recent advances, we propose balanced implementation strategies that maximize innovation while mitigating risks. The review underscores the need for interdisciplinary collaboration to realize the full potential of these technologies in addressing global materials challenges.
The field of materials science has witnessed a transformative shift with the advent of representation learning techniques, particularly for analyzing complex microstructures. This review synthesizes recent conceptual advances in representation learning, including deep neural networks, autoencoders, and vision transformers, applied to microstructure data for tasks such as property prediction, inverse design, and evolution modeling. We explore how these methods extract latent features from high-dimensional microstructure images, enabling efficient computation and discovery of structure-property relationships. However, interpretability remains a significant challenge, as black-box models often obscure the physical meaning of learned representations, hindering trust and scientific insight. We discuss strategies for enhancing interpretability, such as attention mechanisms, heat maps, and post-hoc explanations, drawing from recent studies in alloy microstructures and additive manufacturing. The review highlights the integration of domain knowledge to disentangle representations and address data scarcity issues. By examining case studies in metals, ceramics, and composites, we identify gaps in current approaches, including bias in learned features and limited generalizability across materials classes. Ultimately, this review aims to guide future research toward interpretable representation-learning frameworks that accelerate materials design and foster a deeper understanding of microstructural phenomena.
The ambiguous use of “falsifiability” in materials AI literature poses a significant challenge to the scientific status of AI-generated claims, as researchers frequently present predictive or generative outputs—such as “this perovskite structure is stable at room temperature” or “this inverse-designed alloy exhibits a target bandgap of 1.8 eV”—without clarifying whether these statements could, in principle, be contradicted by empirical observation. Rooted in Karl Popper's philosophy of science and extended through contemporary applications to machine learning, falsifiability serves as the demarcation criterion that distinguishes scientific claims from non-scientific ones by requiring that they logically forbid certain observations rather than merely accommodate data. This paper proposes precise definitions for falsifiable, verified, and testable AI-generated materials claims, tailored specifically to the challenges of data-driven discovery in solid-state systems, generative models, and inverse design. It further introduces a four-component framework for assessing the falsifiability of such claims, centering on claim specification, forbidden observation specification, test design, and falsification protocol. These conceptual foundations carry profound implications for materials AI practice, requiring authors to articulate disconfirming evidence explicitly, reviewers to demand falsifiability statements, and the broader community to adopt standards that elevate predictive modeling from statistical correlation to genuine scientific inquiry. By confronting the boundary between data-driven heuristics and empirically falsifiable science, the present work offers a definitional scaffold that can guide the field toward greater epistemic rigor amid the accelerating integration of artificial intelligence into materials discovery.
In the rapidly advancing field of artificial intelligence for materials science, a persistent and underappreciated limitation has emerged: the overwhelming emphasis on identifying and deploying a single “best” model that maximizes predictive accuracy for properties such as band gaps, formation energies, or mechanical strengths, while largely neglecting the epistemic value of algorithmic diversity across model collections. This paper articulates the theoretical claim that algorithmic diversity functions as a core scientific robustness mechanism, independent of any marginal gains in accuracy, by enabling collective coverage of hypothesis space, resilience to distribution shifts, and more reliable knowledge generation in the face of inherent uncertainties in materials data and modeling assumptions. To operationalize this insight, the work proposes a novel conceptual framework consisting of five interlocking components—diversity dimensions, metrics, generation strategies, robustness linkages, and evaluation protocols—that together redefine how diverse model collections should be designed, assessed, and deployed in materials discovery pipelines. The framework further delineates five distinct types of diversity (architectural, representational, initialization, data-centric, and objective) that each contribute unique robustness benefits when applied to materials-specific challenges such as inverse design or multiscale modeling. By shifting the community’s focus from solitary model optimization to the deliberate cultivation of diverse algorithmic ecosystems, the implications extend to revised authorship practices, peer-review standards, and the establishment of diversity-aware benchmarks, ultimately positioning algorithmic diversity as an essential epistemic virtue for trustworthy, generalizable materials AI.
Scientific path abandonment has emerged as a critical yet underrecognized failure mode in AI-guided materials research, in which promising research directions—such as novel compositional families, structural motifs, or synthesis routes—are terminated prematurely due to insufficient evidence, narrow optimization signals, or algorithmic impatience. This failure mode is defined as the termination of a research direction before sufficient evidence has been gathered to determine its true promise, distinguishing it from rational stopping grounded in conclusive data. The mechanisms driving this abandonment include algorithmic impatience that halts exploration upon short-term metric plateaus, overconfidence in negative predictions, narrow optimization that sacrifices multi-objective potential, and exploration decay inherent in active learning loops. Four distinct types of path abandonment—compositional, structural, synthesis, and property—each generate specific failure modes, such as local optima traps, false-negative cascades, exploration starvation, and regret amplification. Detection principles center on systematic audits, counterfactual reasoning, diversity monitoring, and regret tracking. In contrast, mitigation principles emphasize extended exploration, resource reserves, delayed abandonment thresholds, path revisitation, and regret-aware stopping rules. By articulating this failure mode and offering a comprehensive framework for recognition and remedy, the analysis identifies scientific path abandonment as a systemic risk that undermines the very autonomy and discovery potential that AI promises to deliver in materials science.
In the field of artificial intelligence applied to materials science, a fundamental conflation persists in which exploration noise and scientific error are routinely conflated as interchangeable “mistakes” that must be minimized or eliminated to improve model performance. This paper proposes precise conceptual definitions that separate exploration noise—understood as stochastic variation deliberately or unavoidably introduced into decision-making processes to probe uncertain regions of materials design space—from scientific error, defined as any deviation from ground truth that reduces predictive fidelity, distorts mechanistic understanding, or precipitates incorrect materials decisions without any compensating epistemic gain. The distinction matters profoundly because the systematic elimination of exploration noise eradicates the very mechanism that drives discovery in high-dimensional, data-scarce materials landscapes. In contrast, misclassifying scientific error as mere noise allows systematic flaws to propagate undetected through autonomous discovery pipelines. To resolve this ambiguity, the present work offers a four-criterion framework grounded in intentionality, epistemic benefit, systematicity, and correctability that enables researchers to classify any observed deviation with conceptual clarity. Adoption of this framework carries immediate implications for materials AI practice: it demands new reporting standards that explicitly quantify and justify exploration noise, revised peer-review criteria that interrogate rather than penalize productive randomness, and a cultural shift that reframes stochasticity not as a defect to be denoised but as an essential epistemic resource for accelerating the discovery of novel materials with targeted functionalities.
Materials AI is rapidly converging toward single-model regimes in which a handful of dominant architectures, particularly graph neural networks, have become the de facto standard for property prediction, inverse design, and materials discovery. This model monoculture does not merely reflect technical superiority; it actively produces convergent scientific narratives that shape what the community considers valid knowledge, worthwhile problems, and genuine progress in the field. The present critique identifies four interlocking epistemic risks of this convergence: epistemic narrowing, suppression of alternatives, paradigm lock-in, and the illusion of consensus. These risks threaten the long-term robustness of materials science by limiting the diversity of phenomena that can be observed, the range of methods that can be explored, and the kinds of disagreement that can be productively acknowledged. The consequences include missed discoveries in complex materials systems, methodological stagnation, overconfidence in model outputs, and path-dependent research trajectories that will prove difficult to reverse. Alternative approaches grounded in deliberative methodological pluralism, adversarial benchmarking, narrative diversity, paradigm auditing, and deliberate switching-cost reduction are therefore proposed as necessary correctives if the field is to preserve its epistemic openness while retaining the undeniable benefits of data-driven methods.
Generative models in materials science have emerged as powerful tools for proposing novel atomic structures, compositions, and functional properties. Yet, their scientific evaluation remains conceptually underdeveloped and fragmented across statistical proxies that rarely capture the true relevance to materials. This review systematically examines the conceptual foundations of scientific evaluation for generative materials AI by targeting 30 peer-reviewed publications spanning 2017–2026 and employing a PRISMA-guided methodology focused on evaluation metrics, physical plausibility, chemical validity, synthesizability, novelty, and utility. The evaluation dimensions extend far beyond conventional statistical metrics such as validity percentages or reconstruction error to encompass six interlocking scientific criteria—chemical validity, structural plausibility, property accuracy, synthesizability, novelty, and utility—that together define whether a generated material constitutes a genuine scientific artifact rather than a computational curiosity. Current evaluation practices, as documented across the literature, remain heavily anchored in validity scores, uniqueness counts, and nearest-neighbor novelty checks, with approximately 68% of studies relying primarily on chemical-validity filters and only 22% incorporating any form of synthesizability assessment, revealing a persistent gap between computational convenience and experimental realism. Critical analysis reveals that these practices are necessary yet profoundly insufficient, frequently conflating statistical fidelity with scientific value and overlooking failure modes such as physically unstable geometries or literature-overlooked duplicates. Emerging frameworks, including multi-objective physics-informed scoring, retrospective validation against subsequent experimental discoveries, and downstream task benchmarking, offer promising pathways toward more rigorous standards. Yet significant gaps persist in the absence of community-wide benchmarks, reliable predictors of synthesizability, and domain-specific utility metrics. This review, therefore, offers actionable recommendations for authors, reviewers, and the broader community to elevate generative materials AI from pattern generation to verifiable scientific discovery, ensuring that evaluation protocols align with the epistemological demands of materials science itself.
Artificial intelligence is rapidly moving beyond its early role as a pattern-recognition and predictive-modelling tool in materials science. What began as an acceleration strategy for screening known datasets is now becoming a broader transformation of how materials hypotheses are generated, tested, and refined. The central problem is that this transformation is often described in fragments: predictive models in one literature, generative design in another, physics-informed learning in another, and autonomous laboratories in yet another. A unified conceptual synthesis is needed to explain how these streams collectively move AI from passive assistant to active scientific collaborator. This integrative review traces the evolution of AI in materials science from 2017 to 2026. It frames the field through the idea of the AI co-scientist: an intelligent system that can recognise patterns, propose candidates, incorporate physical constraints, select experiments, and learn from feedback. The review integrates 31 peer-reviewed articles spanning materials informatics, machine learning, generative AI, inverse design, physics-informed modelling, active learning, autonomous experimentation, and self-driving laboratories. It does not present new empirical data, meta-analysis, or bibliometric mapping. The synthesis identifies four major evolutionary stages: pattern recognition, generative design, physics-integrated AI, and autonomous experimentation. These stages are not isolated phases but mutually reinforcing capabilities that increasingly connect computation, synthesis, characterisation, and human judgement. The review concludes that AI is becoming a genuine partner in materials discovery, but this transition depends on trustworthy data infrastructure, interpretable models, robust experimental integration, and new norms for human–AI collaboration. The co-scientist paradigm offers a forward-looking framework for understanding how materials science may be reorganised around closed-loop intelligence.
This review examines the literature on ethical frameworks for artificial intelligence (AI) applied to materials science and discovery, synthesizing insights from 31 peer-reviewed publications spanning 2017 to 2026 to trace the evolution from high-level principles to practical implementation. The methodology involved a systematic search across Web of Science, Scopus, arXiv, and PhilPapers using targeted strings such as “ethics AI materials science,” “responsible AI materials discovery,” “ethical framework AI science,” “dual use materials AI,” “AI ethics principles materials,” “governance AI materials research,” “justice AI materials discovery,” and “value alignment materials AI,” with inclusion limited to peer-reviewed works directly addressing ethical dimensions in scientific or materials contexts, yielding 31 core references after PRISMA-style screening of over 500 initial results. Major ethical principles for AI—beneficence, non-maleficence, autonomy, justice, explicability, and sustainability—are surveyed as foundational guides originally developed in broader AI ethics literature but rarely adapted to materials-specific applications. The current state of materials AI literature reveals a predominant focus on technical acceleration of discovery, with explicit ethical engagement appearing in fewer than 20% of surveyed works and often limited to passing mentions rather than systematic analysis. Materials-specific ethical challenges, including dual-use risks in weaponizable materials, environmental harms from resource-intensive AI-driven synthesis, equity gaps in global access to discoveries, labor displacement through automation, intellectual property ambiguities, and intergenerational justice concerns, remain largely unaddressed despite the field’s rapid growth. Significant gaps persist in operationalizing principles, developing governance mechanisms, and providing domain-tailored guidance, underscoring an urgent need for actionable recommendations to bridge the principles-practices divide and foster responsible materials AI innovation that prioritizes societal benefit, sustainability, and justice.
This review examines the problem of scientific consensus formation in AI-driven materials science by systematically analyzing conceptual approaches from philosophy and sociology of science alongside empirical developments in computational materials research, drawing exclusively on 31 peer-reviewed publications from 2017–2026 identified through targeted searches in Web of Science, Scopus, arXiv, and PhilPapers using terms such as “scientific consensus” AI materials, “consensus formation” machine learning science, “disagreement” materials AI, “benchmark” consensus materials informatics, “epistemic consensus” AI science, “paradigm” materials AI, “scientific disagreement” computational science, and “consensus mechanism” AI research, with inclusion criteria limited to papers addressing epistemology, disagreement, uncertainty, benchmarks, or paradigm dynamics in data-driven disciplines and exclusion of purely technical performance reports. Consensus concepts are traced from logical-positivist agreement on theories through Kuhnian paradigms and Mertonian social processes to Bayesian convergence and pragmatic problem-solving necessities, revealing how each framework illuminates different facets of knowledge coordination in materials science. AI’s impact on consensus formation operates through six distinct mechanisms—accelerated hypothesis validation, model disagreement, benchmark-driven focal points, opacity-induced dissent, data-driven convergence, and authority shifts—both facilitating rapid agreement on material properties and simultaneously generating new forms of epistemic fragmentation. These dynamics create profound tensions and paradoxes, including the trade-off between speed and deliberation, convergence versus diversity, predictive agreement versus explanatory understanding, local versus global consensus, and human versus AI authority, while exposing critical gaps such as the absence of a dedicated theory for AI-mediated consensus, the scarcity of empirical studies tracking real-time consensus processes in materials AI communities, and unresolved questions about managing productive disagreement. Recommendations are offered for researchers, journals, and the broader community to distinguish model agreement from scientific consensus, institutionalize empirical consensus studies, preserve productive dissent, and develop governance protocols that harness AI’s epistemic power without sacrificing critical scrutiny, thereby guiding the field toward more reflexive and robust knowledge production in the age of AI-augmented materials discovery.
Graph-based machine learning has reached maturity for crystalline materials, where periodic boundary conditions and fixed unit cells allow efficient message passing and property prediction. Disordered systems—amorphous solids, liquids, glasses, and polymers—represent the next frontier. These materials dominate everyday technologies, from smartphone screens and optical fibers to electrolytes in batteries and structural alloys, yet they lack the translational symmetry that simplifies graph construction in crystals. This review synthesizes exactly 35 peer-reviewed publications from 2017 to 2025 that apply graph-based learning, graph neural networks (GNNs), and related data-driven methods to non-crystalline materials. The selected works span key target journals including npj Computational Materials, Physical Review B, Journal of Chemical Theory and Computation, Machine Learning: Science and Technology, and Physical Review Materials. The review is organized around a taxonomy of five classes of disordered systems: amorphous solids, liquids, glasses (including supercooled liquids), polymers, and crystals with substitutional disorder. It identifies seven core challenges that distinguish disordered systems from crystals: absence of periodic boundaries, variable graph sizes, heterogeneous local atomic environments, the critical role of medium-range order (5–15 Å), dynamic graph evolution in liquids, lack of standardized benchmarks, and prohibitive computational cost for large simulation boxes. Current methodological approaches are grouped into six categories—local descriptors (non-graph), graph pooling for size invariance, multi-scale GNNs, equivariant architectures, temporal GNNs for liquids, and transfer learning from crystal data—each evaluated for strengths, limitations, and empirical performance on properties such as density, elastic moduli, glass-forming ability, relaxation dynamics, and defect identification. Despite notable progress, three persistent gaps remain: limited size transferability across simulation-box lengths, inadequate capture of medium-range structural correlations, and under-development of methods for dynamic properties. By critically analyzing these 35 studies, this review provides a systematic framework for graph-based learning in disordered materials and highlights open problems that must be solved before these techniques can achieve the same reliability and scalability already demonstrated for crystals. The field stands at an inflection point: continued innovation in graph representations and benchmarking will determine whether graph-based methods can fully unlock predictive modeling for the disordered materials that underpin modern technology.
Extrapolation in disordered materials such as glasses, amorphous solids, and liquids is fundamentally harder than in crystalline systems. Disordered materials lack periodic symmetry, exhibit highly heterogeneous local environments, and possess variable system sizes that range from hundreds to hundreds of thousands of atoms. These characteristics create an exceptionally large hypothesis space for machine learning models and render pure statistical learning approaches unreliable beyond the training distribution. At the same time, traditional physics-based models remain too approximate for quantitative accuracy in complex disordered systems. This conceptual framework proposes a unified model that bridges statistical learning and physics priors to overcome these limitations and enable reliable extrapolation in disordered materials. The framework rests on three core physics priors—locality, smoothness, and invariance—that act as powerful inductive biases. Locality limits interactions to finite cutoffs, smoothness ensures continuous property landscapes, and invariance (rotational, permutation, and size extensivity) dramatically reduces the effective search space. These priors are embedded directly into statistical learning architectures so that the overall prediction combines a physically grounded baseline with data-driven residual corrections. In extrapolation regimes the physics-informed baseline dominates, producing graceful degradation rather than arbitrary outputs. The model further integrates statistical learning components including uncertainty quantification to flag risky predictions, active learning to expand the training distribution adaptively, multi-fidelity strategies that leverage cheap physics approximations, representation learning for cross-system transfer, and ensemble methods for robustness. The resulting conceptual taxonomy clarifies why extrapolation fails in disordered materials and supplies explicit design principles for extrapolation-aware models. This unified approach shifts the default paradigm in computational materials engineering from purely data-driven or purely physics-driven methods toward a hybrid that respects physical constraints while retaining the flexibility of statistical learning. The framework is expected to accelerate discovery in glass design, amorphous polymers, and liquid electrolytes where out-of-distribution generalization is essential. By treating physics priors as non-optional architectural elements rather than optional regularizers, the model offers a practical path toward trustworthy machine learning predictions in the disordered realm.
Lifelong learning, also known as continual or sequential fine-tuning, promises machine-learning models for materials science that continuously improve as new data streams arrive. In theory, a model trained on binary-alloy formation energies should seamlessly incorporate new knowledge about ternary systems, experimental band gaps, or updated DFT databases without losing prior accuracy. Yet in practice, lifelong learning systematically fails for materials property prediction. Sequential fine-tuning triggers catastrophic interference: the model rapidly forgets earlier tasks while attempting to master new ones. This failure mode analysis identifies four core failure modes that are especially destructive in the materials domain. Catastrophic forgetting occurs when shared parameters encoding old task knowledge are overwritten during fine-tuning on new tasks. Negative interference arises when conflicting signals from different data sources force the model into an impossible compromise. Representation overwriting destroys generic feature detectors in early layers that were useful across multiple materials tasks. Task imbalance allows later, larger, or easier datasets to dominate the optimization process, marginalizing earlier knowledge. The root causes are intrinsic to materials science: massive domain shifts between compositional spaces, overlapping but non-identical feature distributions, and the optimizer’s inability to preserve multi-task optima in a shared parameter space. These problems are documented across recent studies of machine-learning interatomic potentials, continual learning in mechanical networks, and sequential task learning on evolving materials databases. This paper provides precise detection signatures for each failure mode and outlines mitigation principles that can be applied immediately by practitioners. By exposing why sequential fine-tuning collapses in materials property prediction, the analysis lays the foundation for genuinely lifelong materials models that accumulate knowledge rather than repeatedly erasing it.
Reproducibility is a cornerstone of scientific integrity, yet in the rapidly expanding field of data-driven materials synthesis, the term is used inconsistently and often misleadingly. Authors frequently claim “reproducibility” without specifying whether they refer to computational verification of models or successful laboratory synthesis of predicted materials. This boundary/Definitional article clarifies that “reproducibility” in this domain encompasses two fundamentally distinct and orthogonal concepts: algorithmic reproducibility and experimental replication. Algorithmic reproducibility is defined as the ability to obtain identical numerical results (within documented floating-point tolerances) when the same code is executed on the same data in the same computational environment. Experimental replication, by contrast, is the ability of an independent laboratory to synthesize the same material—within clearly defined characterization tolerances—by strictly following the published synthesis protocol. The article demonstrates that these two forms of reproducibility operate in separate ontological domains (computational versus chemical) and are frequently conflated in the literature, leading to overclaims, misdirected research effort, and erosion of community trust. Through a systematic boundary analysis, four-level assessment scales are established for each concept, a comparative framework (including an interaction matrix) is presented, and common gray zones and boundary cases are examined. The analysis shows that algorithmic reproducibility is necessary but insufficient for validating materials predictions, while experimental replication is necessary but insufficient for validating the underlying machine-learning pipeline. Only explicit reporting of both at defined levels constitutes a complete and trustworthy claim. This framework provides authors, reviewers, journals, and funders with a precise, operational vocabulary and a practical two-part Reproducibility Declaration standard. Its adoption will reduce ambiguity, strengthen the credibility of AI-driven materials discovery, and accelerate the reliable translation of computational predictions into verifiable laboratory outcomes. The distinctions introduced here are essential for maturing data-driven materials science into a robust, reproducible discipline.
Neural scaling laws describe the power-law decay of prediction error with increasing training dataset size, yet the scaling exponents reported for materials property prediction vary widely (0.1–0.8) across databases and properties. This variation has remained poorly understood, limiting reliable forecasting of data requirements in computational materials discovery. Here we present a purely theoretical framework that attributes the observed differences in scaling exponents primarily to the intrinsic data redundancy of materials databases. We formalize redundancy as the fraction of samples that convey overlapping or duplicate structural information, thereby reducing the effective number of independent samples. Through conceptual derivations and proof sketches grounded in information theory and sample-complexity arguments, we establish that the effective scaling exponent is approximately the ideal (independent-sample) exponent multiplied by (1 − r), where r is the redundancy fraction. Upper and lower bounds on achievable exponents are derived directly from three fundamental structural features: the number of distinct crystal prototypes, elemental compositional diversity, and the variety of local atomic environments. High redundancy—arising naturally from repeated prototypes, compositional biases, and clustered coordination motifs—systematically flattens the exponent and imposes hard structural limits on learning efficiency, independent of model architecture. The framework provides lightweight, model-free metrics for estimating redundancy from standard database statistics and offers actionable guidance for diversity-driven dataset curation. By reframing scaling behavior in terms of effective sample size and structural diversity, this work supplies a unified theoretical explanation for the wide range of reported exponents and a predictive foundation for designing more efficient materials databases.
Foundation models promise a transformative advance in materials science by enabling zero-shot property prediction across diverse chemistries, crystal structures, and physical properties, ostensibly eliminating the need for property-specific labelled datasets and expensive first-principles calculations. Proponents claim that models pre-trained on millions of structures can directly predict formation energies, band gaps, elastic moduli, thermal conductivities, and other attributes for entirely unseen materials without task-specific fine-tuning. This critique demonstrates that such claims largely reflect an illusion of generalizability. Reported zero-shot performance is systematically inflated by six pervasive artefacts: (1) test-set leakage of crystal structures or near-identical analogues from the pre-training corpus, (2) exploitation of strong inter-property correlations, (3) evaluation within the interpolation regime of the pre-training distribution, (4) benchmark bias inherent in widely reused datasets such as the Materials Project, (5) inconsistent pre-processing, data splits, and reporting practices, and (6) the absence of rigorous elementary baselines. Analysis of recent foundation-model literature (2023–2026) shows that zero-shot accuracies frequently collapse once these confounds are controlled and often fail to surpass simple statistical predictors such as k-nearest-neighbour regression on embeddings or linear models based on property correlations. The illusion is particularly acute in materials science due to the repeated reuse of identical structures across property databases and the dense network of physical correlations among computed quantities. Without stringent controls, overstated zero-shot claims risk misdirecting research resources, inflating expectations, and undermining trust in AI-driven materials discovery. This paper provides a diagnostic framework for identifying these artefacts, marshals supporting evidence from the contemporary literature, and proposes a six-point evaluation standard for credible zero-shot claims. Adoption of these practices is essential if foundation models are to deliver genuine advances in out-of-distribution generalization rather than repackaged statistical regularities.
Materials discovery remains painfully slow. Traditional human-driven experimentation, followed by offline machine-learning analysis, requires weeks or months per iteration and leaves vast regions of chemical space unexplored. This position paper argues that autonomous laboratories equipped with real-time ML feedback loops represent not an incremental improvement but a necessary paradigm shift for the future of materials engineering. In these systems, robotic platforms handle synthesis and characterization while ML models continuously update and steer the next experiment, closing the discovery loop in hours rather than weeks. The current paradigm relies on human-in-the-loop decision-making, batch experimentation, and post-hoc ML training. Autonomous laboratories reverse this: robots execute synthesis and characterization tasks, a real-time ML engine analyzes streaming data, and an acquisition function immediately proposes the next candidate, all without human intervention for routine decisions. Early demonstrations have already shown accelerated discovery of battery electrolytes, perovskites, and catalysts. Real-time ML feedback loops demand online learning, rigorous uncertainty quantification, rapid acquisition functions, multi-objective optimization, constraint handling, human oversight for safety, and seamless data streaming. We articulate seven foundational principles for closed-loop discovery: integration-first design, speed as a first-class constraint, uncertainty-driven exploration, graceful degradation, data provenance, modularity, and open standards. These principles address the technical, operational, and cultural barriers that still prevent widespread adoption. While challenges remain—high initial costs, instrument integration, and long-duration experiments—the community now possesses the necessary ML maturity, robotic hardware, and orchestration tools to overcome them. This position calls for coordinated investment in shared autonomous-lab infrastructure, open standards, and training programs so that closed-loop discovery becomes the default workflow across academia and industry. Only then can materials science deliver the energy, sustainability, and electronics breakthroughs society urgently needs.
Foundation models and large language models (LLMs) are rapidly entering materials science, offering new interfaces for property prediction, synthesis planning, and literature mining. This review synthesizes some peer-reviewed publications, focusing on foundation models pre-trained on crystals, molecules, and text, as well as LLM-based approaches for materials tasks. Three primary use cases emerge: (1) property prediction from compositional or structural descriptions, (2) synthesis recipe generation, and (3) extraction of structured data from scientific text. Methods span zero-shot prompting, few-shot prompting, chain-of-thought reasoning, retrieval-augmented generation, full fine-tuning, parameter-efficient fine-tuning, and embedding-based adaptation. What transfers effectively includes generic chemical knowledge, structure–text relationships, qualitative trends, similarity search, and literature extraction. In contrast, what does not transfer includes quantitative property prediction to experimental accuracy, extrapolation beyond pre-training distributions, crystal stability assessment, physics-based reasoning, and hallucination-free synthesis planning. Despite promising demonstrations in common materials, LLMs and foundation models still lag behind specialized graph neural networks in quantitative tasks and fail on compositional or structural novelty. This review provides a systematic taxonomy of applications, a critical analysis of prompting and fine-tuning strategies, and a clear delineation of transfer limitations. Gaps remain in uncertainty quantification, multimodal data scarcity, and rigorous benchmarking against non-LLM baselines. Recommendations for practitioners and developers emphasize realistic expectations and hybrid human–AI workflows to accelerate materials discovery without over-reliance on ungrounded predictions.
Machine learning interatomic potentials (MLIPs) are widely validated against density functional theory (DFT), yet this practice risks reproducing DFT’s systematic biases rather than true experimental behavior. We present a concise four-level hierarchical framework for rigorous experimental validation focused on phonon dispersion and elastic constants. Level 1 evaluates static room-temperature elastic constants; Level 2 assesses zone-center phonon frequencies (Raman/IR); Level 3 examines full phonon dispersion curves from inelastic neutron/X-ray scattering; and Level 4 tests temperature-dependent shifts and softening. For each level we provide key observables, simple comparison metrics, uncertainty-aware criteria, and typical failure modes. A practical validation protocol and minimum reporting standards (Level 1 on ≥3 materials, Level 2 on ≥1) are proposed. This framework complements DFT benchmarks by anchoring MLIPs in experimental reality, improving reliability for thermal, mechanical, and vibrational applications in materials science.
Piezoelectricity is a tensor property of rank 3 that couples mechanical stress to electric polarization in non-centrosymmetric crystals. Accurate prediction of the piezoelectric tensor from atomic structure is essential for discovering new sensors, actuators, and energy-harvesting materials in data-driven materials engineering. Yet many researchers continue to apply invariant graph neural networks such as CGCNN, MEGNet, and SchNet—architectures originally designed for scalar properties like formation energy—to this tensor task. These models are rotationally invariant: their outputs remain unchanged under any reorientation of the crystal lattice. The piezoelectric tensor, however, must transform covariantly with orientation. As a direct consequence, invariant GNNs systematically predict zero piezoelectricity or near-zero random noise across entire datasets. This failure mode analysis identifies four core reasons: sign ambiguity arising from untracked orientation-dependent sign flips, component confusion caused by rotational permutation of tensor elements, symmetry overconstraint that forces unphysical relations across crystal classes, and vanishing off-diagonal components that average to zero under random rotations. The analysis draws exclusively on peer-reviewed literature in computational materials science and shows that equivariant GNNs are not merely an improvement but a strict necessity for any rank-2 or rank-3 tensor property. Detection principles and mitigation strategies are outlined conceptually, without reference to specific benchmarks or numerical scores. The work underscores a broader limitation in invariant architectures for tensorial phenomena in materials informatics and calls for their replacement by orientation-aware equivariant frameworks in future tensor property prediction pipelines.
Generative materials models promise to accelerate discovery by systematically exploring vast regions of chemical space, yet the core concept of “chemical space coverage” remains poorly defined and inconsistently applied across the literature. Researchers routinely claim that their models achieve high “diversity” or “coverage,” but these statements rest on incompatible assumptions about what chemical space actually encompasses. This boundary/definitional article clarifies the term by distinguishing three primary dimensions—compositional space, structural space, and property space—and demonstrates how current diversity metrics conflate or ignore these dimensions, rendering cross-study comparisons unreliable. Drawing on recent advances in generative modelling for crystals and molecules, the analysis shows that a model may report excellent elemental coverage while entirely neglecting novel crystal prototypes or property combinations, or conversely achieve broad structural diversity within a narrow compositional slice. To resolve these ambiguities, the article proposes an operational definition of chemical space coverage built around four explicit, computable metrics: compositional coverage (), structural coverage (), property coverage (), and joint coverage (). Each metric is accompanied by practical boundary conditions that define thresholds for “broad,” “comprehensive,” or “exploratory” coverage. The framework further articulates five essential boundary conditions for sufficiency—task dependence, reference dependence, sparsity adjustment, validity trade-off, and diminishing returns—thereby transforming coverage from a vague aspirational term into a precise evaluative criterion. Adoption of this multi-dimensional framework will enable consistent benchmarking, prevent over-optimistic claims, and guide the responsible development of generative models that truly expand the frontiers of materials design rather than merely resampling known regions. The proposed definitions and boundaries therefore constitute a necessary foundation for the next generation of inverse design methodologies in computational materials engineering.