Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Epistemic Saturation in Materials Informatics: When More Data Stops Adding Meaning
Materials informatics represents a transformative intersection of data science, artificial intelligence, and materials engineering, enabling accelerated discovery and optimization of novel substances through computational analysis. However, this paper introduces the concept of epistemic saturation as a critical threshold where accumulating vast datasets no longer enhances meaningful knowledge generation. Instead, it perpetuates interpretive redundancies and systemic distortions, such as entrenched biases in data curation and algorithmic processing. Drawing on recent advancements in machine learning applications within materials science, we explore the dynamics of data-meaning interactions, highlighting how unchecked scaling of information inputs can lead to diminished epistemic value. The proposed framework interprets these phenomena through feedback structures that reveal trade-offs between quantitative abundance and qualitative insight, emphasizing ethical reasoning in algorithmic design and the need for reflexive systems-level oversight. By synthesizing literature on data integrity, algorithmic limitations, and value-laden scientific practices, this conceptual analysis underscores the implications for sustainable innovation across fields such as alloy development and nanotechnology. Ultimately, recognizing epistemic saturation fosters more integrative approaches to informatics, steering toward resilient knowledge ecosystems that prioritize interpretive depth over mere data proliferation. This shift has the potential to reorient materials research toward epistemically robust outcomes amid the ongoing digital transformation.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 3

Material Spaces Are Not Euclidean: A Conceptual Critique of Distance Metrics in Materials
The conceptualization of material spaces within materials science has traditionally relied on Euclidean distance metrics, yet this approach overlooks the inherent complexities of material properties and structures. This manuscript explores the interpretive dimensions of non-Euclidean geometries in representing material relationships, emphasizing how manifold learning and Riemannian frameworks reveal intricate interaction dynamics among atomic configurations and physical attributes. By synthesizing recent literature on geometric neural operators and hyperbolic embeddings, the analysis underscores the trade-offs between simplified Euclidean assumptions and the richer, curvature-aware interpretations that align with multiscale material behaviors. Conceptual interpretations highlight how distance metrics influence systems-level insights in materials informatics, where flat spaces fail to capture hierarchical or topological nuances. The proposed framework integrates these elements through a steering logic that navigates the epistemic challenges of metric selection, fostering a deeper understanding of material continuity and discontinuity without empirical assertions. Ethical reasoning is woven into considerations of the implications for knowledge representation in computational materials discovery. This critique advocates an integrative view that enhances conceptual coherence in the field by bridging abstract geometric principles with material phenomenology.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 5

Scientific Blind Spots Introduced by Feature Engineering in Materials Informatics
Feature engineering remains central to materials informatics, yet systematically introduces scientific blind spots that constrain discovery and interpretation. These blind spots arise from choices in descriptor selection, transformation, and dimensionality reduction that inadvertently prioritize statistical correlations over physical invariance, overlook multi-scale interactions, and embed dataset-specific biases into model architectures. In small-data regimes common to materials science, engineered features often amplify overfitting while diminishing generalizability across chemical spaces. Interpretability suffers as complex engineered descriptors obscure mechanistic linkages between atomic structure and macroscopic properties. Literature consistently highlights these limitations across perovskites, alloys, energy materials, and porous systems, underscoring the tension between predictive performance and scientific fidelity. This conceptual manuscript synthesizes these challenges and proposes an original Integrated Blind Spot Navigation Model (IBSNM). The framework organizes feature engineering around four interdependent pillars—physical consistency guardrails, multi-scale descriptor integration, uncertainty-aware selection, and iterative co-interpretation—linked by feedback mechanisms that surface and mitigate hidden assumptions. By reframing feature engineering as a navigable landscape rather than a static preprocessing step, the model offers a conceptual pathway toward more robust, transparent materials informatics practices that do not rely on empirical validation.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2023 | Article: 22

Failure Case Reporting in Materials Informatics — What Is Documented and What Is Silenced
Materials informatics, the application of data science and machine learning to materials research, has revolutionized the discovery and design of new materials. However, the field faces significant challenges in reporting failure cases, negative results, and biases, which are often silenced in the literature. This review examines documented failures in materials informatics, such as data bias, model overoptimism, and reproducibility issues, and highlights the systemic factors that lead to their underreporting. Drawing on 30 recent peer-reviewed articles, we explore themes including data quality, algorithmic limitations, and publication bias. The objectives are to assess what is typically documented, identify silenced aspects, such as unsuccessful experiments, and propose strategies for more transparent reporting. By addressing these gaps, the review aims to foster a more robust and trustworthy materials informatics ecosystem, ultimately accelerating sustainable innovation in materials science.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2023 | Article: 29

Recent Advances in Machine Learning-Accelerated Materials Discovery — From Descriptors to Autonomous Experiments
Machine learning (ML) has become a central driver of modern materials discovery, fundamentally reshaping how materials are designed, screened, and experimentally realized. This review examines recent advances in ML-accelerated materials discovery and emphasizes the ongoing progress in material representation and descriptor development toward fully autonomous experimental platforms. We discuss how increasingly sophisticated descriptors—ranging from composition-based features and structure-aware representations to ab initio–derived and learned embeddings—have improved predictive accuracy, data efficiency, and physical interpretability across diverse materials systems. Based on these findings, we discuss the evolution of ML frameworks for property prediction, classification, and inverse design, with particular attention to uncertainty-aware modeling, multiobjective optimization, and explainable learning strategies that bridge predictive performance with scientific insight. The study also highlights the growing role of active learning and generative models in efficiently navigating vast chemical and structural spaces, enabling data-efficient exploration and hypothesis-driven discovery. At the frontier of these developments, autonomous experimental systems integrate ML with robotics to form closed-loop workflows that iteratively design, execute, and refine experiments with minimal human intervention. Applications spanning perovskites, alloys, energy materials, and nanostructures illustrate the broad impact of these approaches in overcoming traditional trial-and-error limitations. Finally, we discuss persistent challenges associated with data scarcity, extrapolation, interpretability, and system integration, and outline future directions toward more robust, scalable, and sustainable autonomous materials discovery. Collectively, these advances represent a paradigm shift from passive data-driven prediction to intelligent, self-guided materials innovation.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2024 | Article: 41

Causal Reasoning in Materials Informatics: A Theory-First Roadmap Beyond Correlation
Materials informatics has achieved rapid progress in predicting composition–structure–property relationships, enabling accelerated screening, surrogate modeling, and exploration of high-dimensional design spaces. However, much of this success remains structurally grounded in correlational learning rather than in explanatory, transportable, or intervention-relevant forms of understanding. This conceptual manuscript argues that correlation-centric models, while often sufficient for ranking candidates under training-like conditions, are epistemically underpowered for high-stakes materials decisions such as processing optimization, microstructural control, deployment certification, and failure-sensitive design, where actions must remain defensible under distribution shift, partial observability, and changing constraints. In such settings, predictive accuracy alone does not establish decision legitimacy: a model may be correct for reasons that do not remain stable under deliberate intervention, confounding, or selection effects, thereby producing actionable recommendations without causal warrant. Motivated by recent developments in structural causal models, causal discovery, counterfactual inference, and invariant representation learning, this paper advances a theory-first reframing: materials AI should be treated as an epistemic instrument whose outputs must be qualified by the causal status they can legitimately support. We propose a novel framework—the Causal Warrant Ladder (CWL)—that classifies materials-model outputs into five ascending levels of causal legitimacy: associative regularities, transportable relations, mechanistic constraints, interventional guidance, and counterfactual design claims. CWL is paired with a Causal-Readiness Map, which specifies the minimal conceptual conditions required for upward movement on the ladder, including identifiability assumptions, invariance structure, intervention semantics, and decision stakes. By separating predictive competence from causal legitimacy, this roadmap provides a disciplined conceptual pathway beyond “black-box correlation” toward materials reasoning that supports robust and responsible design action.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2024 | Article: 49

The Microstructure–Property “Explanation Gap”: A Conceptual Anatomy of Why AI Explanations Often Fail
Artificial intelligence (AI) has become increasingly effective at predicting material properties from microstructure-informed representations, enabling rapid screening and accelerated decision-making. Yet, the “explanations” attached to these predictive systems frequently fail to support the kind of understanding required in microstructure–property science—namely, transferable mechanisms, intervention-relevant guidance, and defensible generalization under realistic shifts in processing, measurement, and operating regimes. This conceptual paper argues that explanation failure in materials AI is often structural rather than incidental: many popular explanation toolkits are optimized for interpreting model behavior rather than for producing scientifically legitimate accounts of why a microstructure yields a property outcome. We define the microstructure–property explanation gap as the persistent mismatch between what explainability tools can formally justify and what materials reasoning demands for action. To anatomize this gap, we identify four recurring causes: representational non-identifiability, confounding by processing history, multi-scale emergence, and instability under distribution shift. Building on this anatomy, we propose a novel theoretical framework—the Explanation Integrity Triad (EIT)—which evaluates any AI explanation along three axes: Representational Integrity, Causal Integrity, and Operational Integrity. The EIT provides a domain-specific vocabulary to prevent mechanistic overclaims and align explanation practices with scientific accountability in applied materials informatics.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2024 | Article: 50

A Theory of Multi-Objective Trade-Offs for Sustainable Materials Optimization with AI
Artificial intelligence (AI) is increasingly positioned as a design partner in materials optimization, enabling accelerated exploration of vast composition–processing–structure spaces under multiple, often conflicting, targets. Yet sustainability-centered materials design is not simply a larger version of multi-property optimization: it requires negotiating trade-offs across heterogeneous objective types such as performance, cost, safety, emissions, toxicity, circularity, and resource criticality, while accounting for lifecycle shifts and stakeholder-dependent priorities. Many current AI-enabled optimization workflows implicitly treat trade-offs as static Pareto-front problems with stable objective meanings and fixed feasibility boundaries. This conceptual manuscript argues that such assumptions are structurally incompatible with sustainable materials decisions, which involve trade-offs that are contextual, value-weighted, and regime-dependent. We introduce a novel theoretical framework—Trade-Off Sensitivity Theory (TOST)—which models sustainability optimization as a decision process governed by objective incompatibility geometry, lifecycle constraint migration, uncertainty-to-consequence coupling, and preference volatility. Rather than proposing algorithms or empirical evaluation, TOST provides a theoretical map linking Pareto efficiency to sustainability legitimacy through three layers: objective semantics, trade-off sensitivity, and action admissibility. The framework clarifies when AI outputs support responsible selection, when optimization is ill-posed, and how sustainable decisions can be justified under conflicting criteria.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 51

Human-in-the-Loop without the Hype: A Conceptual Taxonomy of Human Roles in Materials AI
Human-in-the-loop (HITL) approaches are increasingly invoked in materials artificial intelligence (AI) as a presumed remedy for unreliable models, opaque predictions, and domain-shift failures. Yet “including a human” often functions as a rhetorical assurance rather than a precise scientific claim, masking the fact that humans participate in materially different ways: as labelers, judges, curators, constraint designers, hypothesis framers, risk owners, and accountability anchors. This conceptual manuscript argues that HITL is not a single method but a family of epistemic and governance roles that shape what an AI output means, what it can justify, and what actions it can responsibly warrant. Building on recent developments in materials informatics, active learning, uncertainty quantification, interpretable machine learning, and scientific machine learning, we synthesize a theory-first view of human involvement as a structured intervention in the AI-to-decision pathway rather than an informal override mechanism. We introduce a novel taxonomy that distinguishes (i) where humans intervene in the pipeline (data, representation, model, evaluation, decision), (ii) what kind of authority they exert (epistemic, normative, operational), and (iii) how their involvement changes the legitimacy of downstream claims under differing stakes. The resulting framework replaces HITL hype with a falsifiable conceptual vocabulary for designing responsibility, reliability, and restraint in materials AI.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 52

Scientific Accountability for Materials AI: A Conceptual Standard for Reporting Claims and Limitations
Artificial intelligence (AI) has rapidly expanded the scale and ambition of materials research, enabling property prediction, candidate screening, and data-driven optimization across large chemical and structural spaces. However, the field still lacks a discipline-specific standard for scientific accountability: a structured way to report what an AI output legitimately warrants, under which assumptions, and with what limitations. This gap is not cosmetic; it is epistemic. Materials AI often converts heterogeneous proxies (composition features, crystal graphs, microstructure descriptors) into numerical predictions. Yet, manuscripts frequently present these outputs as claims of generality, mechanism, or design readiness without specifying the scope conditions that would make such claims defensible. Recent progress in graph neural networks, benchmark suites, and large community datasets improves comparability. Still, it also amplifies risks of leakage, distribution shift, and proxy instability, which can inflate conclusions while remaining underreported. Meanwhile, uncertainty quantification and explainable AI are increasingly used as trust signals, even though both can be misunderstood when their semantics are not clearly stated, and their limitations are not operationalized for decision-making. We propose a novel conceptual standard—the Scientific Accountability Sheet (SAS)—which binds reported claims to explicit claim types, scope boundaries, evidence anchors, uncertainty semantics, and decision admissibility. SAS reframes “responsible reporting” as a scientific warrant structure rather than an optional best-practice appendix.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 53

Bias in Materials Datasets without Datasets: A Conceptual Account of How Bias Enters Before Any Modeling
Artificial intelligence (AI) in materials science is often treated as a pipeline in which bias primarily emerges during model training, evaluation, or deployment. This framing is structurally incomplete. Many distortions later labeled as “dataset bias” are already introduced before any dataset is formally assembled, labeled, cleaned, or modeled. This conceptual manuscript advances a theory-first account of pre-dataset bias: systematic misrepresentation that originates upstream of data tables through decisions about what counts as a material instance, a property definition, a valid operating regime, and an actionable target. We argue that early bias is not merely a statistical artifact but an epistemic and procedural commitment that shapes what becomes observable, measurable, and publishable. We introduce a novel framework—the bias before data (BBD) framework—which decomposes pre-dataset bias into five coupled mechanisms: problem framing bias, regime availability bias, measurement–proxy bias, curation–visibility bias, and legitimacy bias. BBD provides a structured vocabulary for identifying where bias enters, why it persists despite technical improvements, and how it constrains the legitimacy of scientific claims even when predictive performance appears strong.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 54

Mechanistic vs. Predictive Success: A Theory of What “Understanding” Means for AI in Materials Science‎
The rise of artificial intelligence (AI) in materials science has highlighted a profound epistemic tension. While AI models excel in predictive accuracy, they often fail to provide mechanistic insights into materials behavior, raising questions about whether such predictions constitute genuine scientific understanding. This tension is particularly acute in materials science, where complex phenomena like phase transitions, defect dynamics, and property emergence demand not only forecasting but also explanatory depth to inform reliable design and innovation. Equating prediction with understanding risks epistemic overreach, potentially leading to unwarranted confidence in AI outputs and hindering progress in fields requiring causal knowledge, such as sustainable materials development. This paper proposes a novel theoretical framework that redefines “understanding” in AI-driven materials research as a multi-layered epistemic construct, distinguishing predictive success from mechanistic insight and actionable knowledge. The framework introduces epistemic validity conditions, interpretive constraints, and decision contexts for evaluating AI contributions, emphasizing alignment with physical principles and the avoidance of semantic inflation. By synthesizing recent literature, it addresses conceptual gaps in current approaches and advocates responsible inference that integrates predictive power with explanatory rigor. This contribution advances philosophical foundations for AI in materials science, fostering more robust, trustworthy scientific practices without empirical validation claims.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 57

Interpretability in Materials AI — What “Explanation” Means and How It Should Be Evaluated Conceptually
The integration of artificial intelligence (AI) and machine learning (ML) into materials science has revolutionized the discovery, design, and optimization of new materials, enabling accelerated predictions of properties and behaviors previously unattainable with traditional methods. However, the “black-box” nature of many advanced AI models poses significant challenges, including a lack of transparency that hinders scientific understanding, trust, and practical adoption in materials research. This narrative review explores the concept of interpretability in materials AI, focusing on what constitutes an “explanation” and how it should be conceptually evaluated. Drawing from recent advancements in explainable AI (XAI), we delineate definitions of explanations tailored to materials informatics, emphasizing their role in bridging computational predictions with physical insights. We examine thematic aspects such as intrinsic versus post-hoc interpretability methods, the multidimensional nature of explanations (e.g., local vs. global, feature-based vs. mechanistic), and conceptual frameworks for evaluation, including criteria like fidelity, comprehensibility, robustness, and domain-specific relevance. By synthesizing the literature, we highlight how explanations can enhance materials discovery across alloy design, catalyst development, and polymer engineering, while addressing gaps in current evaluation practices. The review underscores the need for standardized conceptual metrics that go beyond quantitative benchmarks to incorporate qualitative, human-centered assessments in materials science contexts. Ultimately, this work aims to guide researchers toward developing interpretable AI systems that not only predict but also elucidate underlying material phenomena, fostering a more insightful and ethical application of AI in materials innovation.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2024 | Article: 64

Causality in Materials Informatics — Conceptual Progress, Limitations, and Future Directions
Materials informatics has emerged as a central paradigm in contemporary materials science, leveraging machine learning and data-driven modeling to accelerate materials discovery, optimization, and deployment. Despite substantial advances in predictive accuracy, most existing approaches remain fundamentally correlational, limiting their reliability under distribution shifts, experimental interventions, and real-world deployment scenarios. This reliance on correlation constrains scientific interpretability and undermines the capacity of AI systems to function as genuine instruments of materials reasoning. Causality offers a principled framework for overcoming these limitations by explicitly modeling cause-and-effect relationships among composition, processing, structure, and properties. This narrative review synthesizes conceptual progress in integrating causal inference into materials informatics, examining foundational causal frameworks, advances in causal discovery, and hybrid causal–machine learning approaches, and emerging applications across materials domains such as nanocatalysis, ferroelectrics, and electrochemical energy storage. We critically analyze persistent challenges—including data scarcity, assumption violations, limited external validity, and computational and epistemic constraints—that currently hinder widespread adoption. Drawing exclusively on peer-reviewed literature published, the review emphasizes thematic and epistemic developments rather than algorithmic prescriptions. We argue that causality represents a structural shift in how AI systems contribute to materials science: from correlational predictors to intervention-aware, mechanism-aligned reasoning tools. By articulating future directions centered on hybrid modeling, domain-knowledge integration, and interdisciplinary collaboration, this review positions causality as a necessary foundation for robust, generalizable, and scientifically legitimate materials informatics.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2025 | Article: 66

Boundary Conditions of Transfer Learning in Materials Science: A Conceptual Theory of When Knowledge Transfers Fail
Transfer learning has emerged as a pivotal strategy in materials science, enabling the reuse of knowledge from data-rich domains to inform predictions in data-scarce contexts, thereby accelerating discovery across alloy design, nanomaterials, and functional compounds. Despite its growing adoption, the effectiveness of transfer learning remains contingent on subtle boundary conditions that delineate productive knowledge integration from ineffective or counterproductive transfer. This conceptual paper develops a theoretical framework to interpret these boundaries by examining interaction dynamics between source and target domains in materials contexts. It explores how mismatches in representational hierarchies—such as between atomic-scale and macroscopic descriptions—disrupt knowledge flow and yield distorted predictive outcomes. Systems-level analysis reveals trade-offs in model adaptability, where reliance on pre-trained representations may obscure emergent properties specific to target materials. Ethical considerations further highlight the risks of bias propagation from simulated to experimental domains, with implications for research prioritization and resource allocation. By integrating perspectives from materials informatics and complexity theory, the framework articulates steering logics to mitigate transfer failures through adaptive feature alignment. This work advances conceptual understanding of transfer learning limitations and provides interpretive guidance for future AI integration in materials science, without empirical validation.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2025 | Article: 79

Data Is Not Neutral: A Conceptual Framework for Value-Laden Measurement Choices in Materials Informatics
Materials informatics has become a central paradigm in materials science, leveraging machine learning and large-scale datasets to accelerate property prediction, discovery, and design. However, prevailing approaches often treat data as a neutral substrate for modeling, obscuring the value-laden processes through which data is generated. Measurement choices—what properties to quantify, which materials to prioritize, and which experimental or computational protocols to employ—are inherently shaped by epistemic commitments, practical constraints, and broader societal priorities. These choices embed values into data infrastructures, systematically influencing which material phenomena become visible and which remain obscured in downstream models. This manuscript advances a conceptual framework that interprets measurement choices as value-mediated interfaces linking scientific priorities to data constitution and modeling feedback in materials informatics. The framework elucidates how value horizons, choice architectures, data formation processes, and modeling circuits interact to produce steering logics, trade-offs, and path-dependent dynamics. By reframing data bias as a constitutive outcome of value-conditioned measurement rather than a purely technical artifact, the framework reveals characteristic failure modes—including value lock-in, patterned absences, and self-reinforcing feedback—that constrain epistemic exploration. Integrating insights from materials informatics, data bias studies, and philosophical analyses of scientific practice, the framework provides a diagnostic lens for understanding the non-neutrality of data in iterative AI-driven workflows. Rather than prescribing methodological interventions, it foregrounds the epistemic consequences of measurement decisions, inviting greater reflexivity in shaping data landscapes over time. This perspective repositions materials informatics as an evolving epistemic system whose possibilities and limits are co-produced by values, measurements, and models.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2025 | Article: 84

From Correlations to Design Rules: A Conceptual Model of Knowledge Extraction in Materials AI
The integration of artificial intelligence (AI) into materials science has substantially accelerated property prediction and materials screening. Yet, the predominance of data-driven correlations has exposed a persistent epistemic gap between predictive success and the derivation of interpretable, generalizable design rules. This conceptual manuscript develops a theoretical framework for knowledge extraction in materials AI that explicitly addresses this gap by reframing the transition from correlations to design rules as a staged epistemic process rather than a by-product of model performance. Drawing on literature in materials informatics, data bias, and philosophy of science, the framework organizes knowledge extraction into four interconnected stages—Correlation Mapping, Bias Interrogation, Value Integration, and Rule Synthesis—linked through continuous epistemic validation. The model foregrounds epistemic agency, requiring explicit scrutiny of assumptions, biases, and value commitments before causal inference. Six propositions articulate the conditions under which AI-derived correlations may legitimately support prescriptive design claims, emphasizing reflexive feedback and epistemic governance. By conceptualizing knowledge extraction as a norm-governed process of justification, this work provides a theoretical scaffold for transforming AI outputs into scientifically defensible design rules, contributing to a more reliable and responsible epistemology of materials discovery.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 85

Human Oversight as System Design: A Conceptual Reframing of Control in Semi-Autonomous Materials AI
The integration of artificial intelligence (AI) into materials science has ushered in an era of semi-autonomous systems that accelerate discovery through predictive modeling, high-throughput screening, and adaptive experimentation. These systems offer substantial promise for addressing global challenges in energy, sustainability, and advanced manufacturing; however, their reliance on data-driven inference introduces risks related to bias propagation, epistemic uncertainty, and misalignment with scientific values. Conventional approaches treat human oversight primarily as an external corrective mechanism—post hoc monitoring or intervention in response to model outputs. This paper proposes a conceptual reframing wherein human oversight is repositioned as an intrinsic element of system design. Rather than viewing control as supervision layered atop an autonomous core, oversight is conceptualized as deliberate architectural choices that embed human judgment into the foundational structure of semi-autonomous materials AI. Drawing on literature from materials informatics, data bias mitigation, explainable AI, and human-AI collaboration, the proposed framework delineates three interdependent dimensions: epistemic boundary-setting, value-aligned modulation, and adaptive reflexivity. This reframing shifts the discourse from mitigating human absence to engineering human presence, fostering systems that are inherently more robust, interpretable, and aligned with the normative goals of scientific inquiry. By reconceptualizing oversight as design, the framework offers a pathway to responsible integration of AI in materials discovery without presupposing full autonomy or diminishing human agency.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 86

The Limits of Scale in Materials AI: A Conceptual Argument Against ‘Bigger Is Always Better’
In the rapidly evolving field of materials artificial intelligence (AI), the prevailing emphasis on scaling data volumes and computational resources has driven significant advancements in predictive modeling and discovery processes. However, this conceptual manuscript interrogates the implicit assumption that larger scales invariably yield superior outcomes, positing instead that unchecked expansion introduces intricate interaction dynamics that undermine the integrity of materials informatics. Through an integrative analysis, we explore how escalating data scales interact with inherent biases, leading to amplified distortions in representational fidelity and epistemic reliability. The framework delineates trade-offs wherein quantitative abundance may erode qualitative depth, fostering feedback structures that perpetuate homogeneity in material explorations at the expense of diversity. Ethical reasoning underscores the epistemic implications, revealing how scale-driven approaches can inadvertently prioritize dominant paradigms, marginalizing underrepresented material classes and contexts. Systems-level insights highlight steering logics that balance scale with interpretive nuance, advocating for calibrated integrations that preserve domain-specific insights. This argument reframes scale not as an unequivocal virtue but as a contingent factor within broader conceptual interpretations, urging a reevaluation of priorities in applied AI for materials science to foster sustainable and equitable progress.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 87

Uncertainty-Conditioned Experiment Planning: A Conceptual Framework for AI-Guided Materials Exploration
Materials exploration faces persistent challenges stemming from vast chemical spaces, high experimental costs, and inherent uncertainties in predictive models. While machine learning has accelerated property prediction and guided candidate selection, conventional approaches often treat uncertainty as a uniform metric within fixed acquisition strategies. This conceptual paper introduces uncertainty-conditioned experiment planning (UCEP) as a novel theoretical framework for AI-guided materials discovery. UCEP reframes experiment planning as a dynamic process conditioned on the multidimensional character of uncertainty, integrating epistemic and aleatoric components, data-related biases, and model limitations into the steering logic. Rather than relying on static acquisition functions, the framework emphasizes adaptive interaction dynamics between uncertainty characterization and planning decisions, enabling context-sensitive trade-offs between exploration, exploitation, and bias mitigation. Drawing on interpretive insights from materials informatics and uncertainty quantification literature, UCEP highlights systems-level feedback structures that can enhance epistemic robustness and scientific efficiency without presupposing empirical outcomes. The framework offers analytical implications for rethinking how AI systems interpret and respond to uncertainty in iterative discovery cycles, contributing to more reflective and integrative AI-assisted materials research.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 88

Scientific Claims Under Distribution Shift: A Conceptual Theory for Robust Materials AI Inference
The integration of artificial intelligence into materials science has accelerated property prediction, inverse design, and discovery pipelines. Yet, the reliability of resulting scientific claims remains vulnerable to distribution shifts—systematic differences between training and inference data distributions arising from variations in synthesis protocols, characterization instruments, environmental conditions, or sampling biases. This purely conceptual manuscript develops a novel theoretical framework for robust materials AI inference in the presence of such shifts. We posit that distribution shifts do not merely degrade predictive accuracy but fundamentally alter the epistemic status of scientific claims by introducing unaccounted covariances between material descriptors and latent generative processes. The framework reconceptualizes inference as a multi-layered epistemic process: (i) shift ontology delineation, (ii) value-laden alignment of data representations with domain invariants, and (iii) claim robustness via counterfactual stabilization. By synthesizing insights from materials informatics, machine learning theory on distribution shifts, and philosophical analyses of epistemic values in science, we argue that robust inference requires explicit modeling of shift-induced epistemic uncertainty rather than mitigation as a post hoc engineering concern. This theory provides a conceptual scaffold for evaluating the validity of AI-derived materials claims across heterogeneous datasets, advancing a shift from performance-centric to epistemically grounded AI deployment in materials science.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 89

Small-Data and Sparse-Regime Learning in Materials AI — Methods, Assumptions, and Limits
The integration of artificial intelligence (AI) and machine learning (ML) into materials science, often referred to as materials informatics or materials AI, has accelerated the discovery, design, and optimization of advanced materials. However, materials science frequently operates in small-data and sparse-regime conditions, where datasets are limited in size (often tens to hundreds of samples), high-dimensional, imbalanced, or sparsely populated due to the high cost, time, and complexity of experimental measurements and high-fidelity simulations. This narrative review synthesizes recent advances in methods tailored to these constraints, categorizing approaches at the data-source level (e.g., literature extraction, database construction, high-throughput workflows), algorithmic level (e.g., support vector machines, Gaussian process regression, ensemble models, imbalanced learning techniques), and strategic level (e.g., active learning, transfer learning). Key assumptions underlying these methods are examined, including similarity between source and target domains for transfer learning, representativeness of initial samples and reliable uncertainty quantification in active learning, and the validity of physical priors or inductive biases in physics-informed approaches. The review also addresses inherent limits, such as risks of overfitting, poor generalization beyond the training distribution, sensitivity to data quality and noise, challenges in uncertainty calibration, and dependence on domain expertise. By highlighting successful applications in property prediction, alloy design, and perovskite optimization, this work elucidates the current capabilities and boundaries of small-data and sparse-regime learning in materials AI, guiding researchers navigating data-limited environments.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2026 | Article: 90

Physics-Integrated Machine Learning for Materials Science — Conceptual Taxonomies and Open Questions
The integration of physical principles into machine learning (ML) frameworks has emerged as a transformative approach in materials science, addressing the limitations of purely data-driven models by incorporating domain knowledge to enhance predictive accuracy, generalizability, and interpretability. This narrative review explores the conceptual taxonomies of physics-integrated ML methods, their applications in materials discovery and design, and the associated challenges in data bias and ethical considerations. Drawing on recent peer-reviewed literature, we classify physics-integration strategies such as physics-informed neural networks (PINNs), hybrid models combining ML with physical simulations, and constraint-based learning, and highlight their roles in solving complex problems such as material property prediction, microstructure analysis, and phase stability. We also examine how data biases in training datasets can propagate errors and inequities in model outputs, and discuss the ethical values underpinning the use of AI in scientific research, including transparency, accountability, and societal impact. The review underscores the potential of these methods to accelerate innovation in materials science while emphasizing the need for rigorous validation and interdisciplinary collaboration. By synthesizing current advancements, this article aims to provide a foundational understanding for researchers and practitioners, paving the way for future developments in this interdisciplinary field.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2026 | Article: 91

Conceptual Foundations for Adversarial Validation in Materials Machine Learning
Standard validation protocols in materials machine learning continue to rely on the assumption that training and test data are drawn from the same underlying distribution. This assumption is almost invariably violated in real-world materials datasets because of temporal drift in measurement techniques, compositional biases in database construction, and experimental confounders arising from different laboratories and instruments. This conceptual framework article proposes adversarial validation as a diagnostic tool specifically tailored for materials informatics: a method that trains a discriminator to explicitly detect whether a distribution shift exists between any two datasets, thereby revealing hidden generalization failures that conventional train-test splits and k-fold cross-validation cannot expose. The framework introduces the conceptual foundations of adversarial validation, distinguishes it from adversarial attacks, articulates why the technique is particularly powerful in the small-data, high-dimensional, and physically constrained domain of materials science, and offers a five-component structure for its systematic application—feature-space definition, classifier selection, shift-detection thresholding, localization of driving features, and actionable response rules. By embedding materials-specific domain knowledge into the interpretation of discriminator performance, the approach transforms validation from a passive checkpoint into an active diagnostic that can distinguish temporal shift from compositional bias and experimental confounding. The implications for materials AI practice are immediate and transformative: researchers can now report adversarial validation results alongside standard metrics, trigger targeted dataset augmentation or model retraining when shifts are detected, and document potential sources of distribution mismatch in experimental workflows, ultimately raising the robustness and trustworthiness of property predictions that underpin materials discovery and design.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 99

Algorithmic Forgetting as a Design Choice: A Conceptual Analysis of Memory in Materials AI
The term “forgetting” appears throughout the materials artificial intelligence literature in multiple, often contradictory senses: as a catastrophic failure that destroys previously acquired knowledge of structure–property relations, as an unexamined side effect of data deletion or replay buffer limits, and occasionally as an implicit consequence of model capacity constraints. This conceptual ambiguity impedes precise communication, obscures design decisions, and prevents the field from treating forgetting as a controllable parameter rather than an inevitable defect. The present boundary/definitional paper proposes a precise definition of algorithmic forgetting as a deliberate design choice, distinct from both catastrophic forgetting and passive capacity limits. It distinguishes algorithmic forgetting from five nearby concepts—catastrophic forgetting, data deletion, privacy preservation, capacity saturation, and regularization-induced compression—by clarifying intent, mechanism, epistemic consequences, and reversibility. The paper further articulates the conditions under which forgetting becomes beneficial (adaptation to distribution shift in experimental data streams, selective retention under resource constraints, and controlled deletion for intellectual property or safety) versus harmful (loss of rare but physically valid examples). Finally, it supplies a materials-specific conceptual framework for deciding what to forget and what to retain, grounded in rarity, recency of validation, and relevance to the current search space. By reframing forgetting as an explicit design lever, this analysis offers materials AI practitioners a shared vocabulary and a systematic approach to engineering memory policies that enhance rather than undermine long-term scientific utility.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 100

The Measurement Problem in Materials Informatics: When Observing Changes in the System
In materials informatics, the act of measuring a material property is routinely treated as a neutral act of passive observation. Yet, every measurement consumes finite resources, physically alters the sample, or reshapes the space of future measurements through model-guided selection. This paper identifies a direct analog of the quantum measurement problem within data-driven materials discovery: observation is not merely informative but constitutively changes the system being observed by depleting experimental budgets, inducing material modifications, and biasing the very distribution of data that subsequent AI models will learn. The theoretical claim advanced here is that materials informatics harbors an intrinsic measurement problem in which AI-guided measurement actively constructs rather than neutrally samples the observable landscape, thereby rendering the resulting datasets and models path-dependent on the history of prior observations. Key concepts include resource depletion, selection feedback loops, and measurement-driven evolution, all of which distinguish classical materials measurement effects from quantum collapse while sharing the core epistemic feature of non-neutrality. The implications are far-reaching for AI-guided materials discovery: autonomous laboratories must treat measurement policies as interventions rather than recordings, active-learning algorithms must internalize the cost of altering the observable world, and dataset curation protocols must document measurement history as rigorously as they document final property values. By theorizing this measurement problem, the present analysis offers a conceptual framework that reframes experiment design, model training, and discovery workflows as inherently self-referential processes in which the observer and the observed co-evolve.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 101

Model Entropy and Scientific Information Loss in Compressed Representations of Materials
Compressed representations—such as handcrafted descriptors, autoencoder embeddings, and graph-neural-network latent spaces—have become indispensable in artificial-intelligence-driven materials science because they enable scalable property prediction from high-dimensional atomic configurations. Yet the very act of compression, while optimizing statistical correlation with target properties, systematically discards information whose scientific value lies outside mere predictive utility. This theoretical analysis applies information-theoretic principles from Shannon and Cover and Thomas to examine how dimensionality reduction in materials representations affects the retention of scientifically relevant content. Drawing on the concept of model entropy introduced by S. S., the paper introduces “model entropy” as a quantitative lens for assessing the information content preserved in any compressed materials representation. It articulates a core theoretical claim: compression optimized for predictive accuracy maximizes statistical information but can erode scientific information—mechanistic, causal, and counterfactual structures essential for understanding, explanation, and extrapolation. A typology of five distinct information-loss mechanisms is developed, each illustrated with representative materials-science scenarios. The analysis culminates in concrete implications for representation design and scientific inference, arguing that future materials AI must move beyond accuracy-centric evaluation toward explicit auditing and preservation of scientific information. By distinguishing statistical signal from epistemic content, this work offers a conceptual framework for building representations that serve both prediction and discovery without hidden epistemic costs.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2023 | Article: 112

The Handling of Domain Shift in Materials Machine Learning Literature: A Review Study
This review systematically examines the handling—or more often the neglect—of domain shift within the materials machine learning literature published between 2017 and 2023, drawing on a targeted search of peer-reviewed publications across specialized databases and journals to compile and analyze exactly 30 representative studies that span foundational overviews, application-focused works, and methodological explorations. Domain shift in materials science takes four distinct yet interrelated forms—temporal, compositional, experimental, and theoretical—each arising from the inherently heterogeneous nature of materials data sources that range from evolving laboratory protocols and diverse chemical families to inter-laboratory variations and discrepancies between computational approximations and experimental realities. Current practices reveal that explicit acknowledgment of domain shift remains rare, with the majority of papers proceeding under the default assumption of identical training and test distributions. At the same time, detection methods and adaptation strategies appear in fewer than one in five studies, leaving models vulnerable to silent degradation when deployed on real-world materials problems. The surveyed methods for handling domain shift include statistical detection techniques, domain-adversarial training frameworks, feature-alignment approaches, and shift-robust evaluation protocols, many of which have been proposed in adjacent machine-learning fields yet remain underutilized in materials contexts despite their direct relevance to property prediction and inverse design tasks. Collectively, these findings underscore the urgent need for standardized shift-reporting protocols, the development of materials-specific out-of-distribution benchmarks, and the integration of domain-adaptation pipelines into routine workflows, thereby elevating the reliability, generalizability, and practical utility of machine-learning models in accelerating materials discovery.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2023 | Article: 117

A Conceptual Theory of Measurement Validity for AI-Generated Materials Properties
The pervasive reliance on predictive accuracy metrics such as mean absolute error, root mean square error, and R² in materials artificial intelligence has created a fundamental misconception: that low prediction error equates to a valid measurement of a material’s property. This paper argues that accuracy alone is insufficient because an AI-generated property value may align closely with held-out test data yet fail to support the specific scientific or engineering inferences for which it is intended. Drawing on foundational measurement validity theory from psychometrics and the social sciences, the manuscript adapts these concepts to the unique context of AI-generated materials properties. It proposes a novel five-component conceptual theory of measurement validity tailored to machine-learning predictions of physical quantities such as band gaps, formation energies, and mechanical moduli. Five distinct dimensions of validity—construct, criterion, generalizability, robustness, and consequential—are articulated and illustrated with materials-specific scenarios. Finally, the framework offers concrete implications for authors, reviewers, and the broader materials informatics community, shifting validation practices from narrow accuracy reporting toward comprehensive evidence-based arguments that link predictions to intended uses. By distinguishing accuracy from validity, this conceptual framework aims to elevate the epistemological rigor of AI-driven materials discovery and design.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2024 | Article: 119

Conceptual Treatments of Causality in Materials Informatics — From Correlation to Intervention: A Review Study
This review systematically examines the conceptual treatment of causality within materials informatics literature published between 2017 and 2024, drawing exclusively on a curated set of 26 studies identified through targeted and broadened searches across Web of Science, Scopus, arXiv, and specialized databases using terms such as “causal inference,” “causality materials informatics,” “structural causal model,” “directed acyclic graph,” “intervention materials design,” and “counterfactual materials prediction,” with inclusion criteria focused on relevance to materials AI while allowing broader engineering and general causal frameworks where they intersect with materials problems. The analysis reveals a pronounced dominance of correlation-based approaches in materials artificial intelligence, where predictive models achieve impressive statistical fits for structure-property relationships yet seldom progress to robust causal claims, as evidenced by the majority of surveyed works prioritizing accuracy metrics over interventional or counterfactual reasoning. Key causal concepts and frameworks, primarily drawn from Pearl’s foundational hierarchy of association, intervention, and counterfactuals as well as structural causal models and directed acyclic graphs, are introduced and contrasted with their limited adoption in the field. Causal methods that have been applied, albeit sparingly, to materials informatics—ranging from data-driven causal discovery to Bayesian causal modeling—are surveyed alongside their strengths and context-specific limitations. Persistent challenges, including the rarity of randomized interventions in experimental materials workflows and the confounding effects inherent in high-dimensional observational datasets, are highlighted as barriers that leave substantial gaps in the literature. Ultimately, this review offers targeted recommendations for authors, reviewers, and the broader community to integrate causal reasoning more explicitly, thereby moving materials informatics from correlational prediction toward actionable intervention and counterfactual understanding essential for autonomous materials design.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2024 | Article: 126
Filters
Clear All





Access type