Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Inference Without Ground Truth: A Conceptual Theory of Validation in Materials AI
In the domain of materials artificial intelligence (AI), the lack of reliable ground truth poses significant challenges for validating inferential processes. This conceptual manuscript develops a novel theoretical framework for understanding validation dynamics in contexts where empirical benchmarks are scarce or contested. Drawing on recent literature in materials informatics, data bias, and epistemic values in science, the framework interprets validation as an integrative system of interaction dynamics between AI-generated inferences and epistemic feedback structures. It explores the analytical implications of managing trade-offs between uncertainty and bias, emphasizing systems-level insights into how inferential reliability emerges from iterative conceptual interpretations rather than direct empirical confrontation. The framework highlights ethical reasoning in steering logics that govern data curation and model deployment in materials discovery. By synthesizing these elements, the paper offers interpretive tools for navigating the epistemic landscape of AI-driven materials science, fostering more robust conceptual integration without relying on propositional claims or empirical validation. This approach contributes to applied AI in materials by illuminating pathways for enhanced inferential integrity amid inherent data ambiguities.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 1

The Coordination Problem in Multi-Model Materials AI Pipelines
Materials science increasingly relies on artificial intelligence (AI) pipelines that integrate multiple models of varying fidelities, architectures, and objectives to accelerate discovery and design. These multi-model workflows—encompassing low-fidelity approximations, high-fidelity simulations, machine learning surrogates, and experimental feedback—promise efficiency but introduce a fundamental coordination problem: reconciling disparate predictions, managing conflicts, ensuring interoperability, and mitigating emergent behaviors or bottlenecks. This conceptual manuscript examines the coordination challenges in such pipelines, drawing on recent advances in multi-fidelity learning, active learning, hybrid modeling, and workflow orchestration. It analyzes how integration conflicts arise from differences in scale, accuracy, and data provenance, potentially leading to consensus failures, validation cascades, and optimization bottlenecks. The discussion highlights conceptual strategies for robust coordination, including uncertainty-aware fusion, adaptive sampling, and iterative refinement, while underscoring the need for principled frameworks to harness the full potential of multi-model systems in materials AI.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 2

Epistemic Saturation in Materials Informatics: When More Data Stops Adding Meaning
Materials informatics represents a transformative intersection of data science, artificial intelligence, and materials engineering, enabling accelerated discovery and optimization of novel substances through computational analysis. However, this paper introduces the concept of epistemic saturation as a critical threshold where accumulating vast datasets no longer enhances meaningful knowledge generation. Instead, it perpetuates interpretive redundancies and systemic distortions, such as entrenched biases in data curation and algorithmic processing. Drawing on recent advancements in machine learning applications within materials science, we explore the dynamics of data-meaning interactions, highlighting how unchecked scaling of information inputs can lead to diminished epistemic value. The proposed framework interprets these phenomena through feedback structures that reveal trade-offs between quantitative abundance and qualitative insight, emphasizing ethical reasoning in algorithmic design and the need for reflexive systems-level oversight. By synthesizing literature on data integrity, algorithmic limitations, and value-laden scientific practices, this conceptual analysis underscores the implications for sustainable innovation across fields such as alloy development and nanotechnology. Ultimately, recognizing epistemic saturation fosters more integrative approaches to informatics, steering toward resilient knowledge ecosystems that prioritize interpretive depth over mere data proliferation. This shift has the potential to reorient materials research toward epistemically robust outcomes amid the ongoing digital transformation.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 3

Algorithmic Attention as Scientific Bias: A Conceptual Analysis for Materials AI
The rapid integration of machine learning and artificial intelligence into materials science has introduced powerful capabilities for predicting, screening, and discovering new materials. Yet this integration also engenders a distinctive form of bias that operates not merely through skewed training data but through the mechanisms by which models allocate and distribute attention across chemical, structural, and property spaces. This paper conceptualizes “algorithmic attention” as a form of scientific bias that manifests in materials AI systems, shaping which phenomena receive emphasis, which regions of materials space are explored, and ultimately which knowledge claims gain epistemic legitimacy within the field. Attention is interpreted here as the patterned prioritization embedded in model architectures, loss functions, data sampling strategies, and iterative feedback loops between prediction and experiment. The analysis explores how such attention dynamics amplify existing data imbalances, create self-reinforcing discovery loops, misalign interpretive authority between model outputs and domain expertise, complicate validation of uncertain predictions, steer research trajectories through hidden optimization priorities, and pose system-level challenges for epistemic reliability and governance. Drawing on recent literature in materials informatics, bias in machine learning, and philosophy of data-driven science, the paper develops an integrative conceptual framework that treats algorithmic attention as an emergent property of socio-technical knowledge systems rather than a purely technical artifact. This framing highlights trade-offs between predictive scalability and epistemic pluralism, underscoring the need for reflective practices that render attention mechanisms more visible and contestable within materials discovery workflows.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 4

Material Spaces Are Not Euclidean: A Conceptual Critique of Distance Metrics in Materials
The conceptualization of material spaces within materials science has traditionally relied on Euclidean distance metrics, yet this approach overlooks the inherent complexities of material properties and structures. This manuscript explores the interpretive dimensions of non-Euclidean geometries in representing material relationships, emphasizing how manifold learning and Riemannian frameworks reveal intricate interaction dynamics among atomic configurations and physical attributes. By synthesizing recent literature on geometric neural operators and hyperbolic embeddings, the analysis underscores the trade-offs between simplified Euclidean assumptions and the richer, curvature-aware interpretations that align with multiscale material behaviors. Conceptual interpretations highlight how distance metrics influence systems-level insights in materials informatics, where flat spaces fail to capture hierarchical or topological nuances. The proposed framework integrates these elements through a steering logic that navigates the epistemic challenges of metric selection, fostering a deeper understanding of material continuity and discontinuity without empirical assertions. Ethical reasoning is woven into considerations of the implications for knowledge representation in computational materials discovery. This critique advocates an integrative view that enhances conceptual coherence in the field by bridging abstract geometric principles with material phenomenology.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 5

Synthetic Data as Scientific Intervention: A Conceptual Framework for Materials AI
In the evolving landscape of materials artificial intelligence (AI), synthetic data emerges not merely as a technical augmentation but as a profound scientific intervention that reshapes the interpretive dynamics of knowledge generation. This manuscript develops a conceptual framework that interprets synthetic data as an intermediary layer facilitating interactions between empirical realities and algorithmic abstractions in materials science. This study synthesizes recent literature and examines how synthetic data influences epistemic trade-offs, such as those between data fidelity and model generalizability. It steers feedback structures within AI-driven discovery processes. The framework underscores systems-level insights into integrating generative models with domain-specific ontologies, highlighting ethical considerations in the curation of virtual datasets that mirror physical constraints without empirical grounding. Analytically, it explores the implications for accelerating materials innovation through enhanced representational capacities, while addressing potential distortions in scientific reasoning arising from over-reliance on simulated inputs. This interpretive approach reveals the transformative potential of synthetic data in reconfiguring the boundaries of human-AI collaboration, fostering a more reflexive understanding of material phenomena. Ultimately, the framework invites a reevaluation of data’s role in scientific inquiry, emphasizing integrative logics that balance innovation with epistemological integrity in the pursuit of advanced materials.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 6

The Role of Surprise in AI-Driven Materials Discovery
The integration of artificial intelligence (AI) into materials discovery processes introduces dynamic elements that reshape traditional paradigms of scientific inquiry. This manuscript explores the conceptual role of surprise—understood as unexpected deviations in predictive models or exploratory outcomes—within AI-driven frameworks for identifying novel materials. Through an interpretive lens, it examines how surprise serves as a steering mechanism in iterative learning cycles, influencing the balance between exploiting known material properties and exploring uncharted compositional spaces. The synthesis of recent literature highlights emergent patterns in which AI systems, by encountering anomalous data or unanticipated correlations, facilitate shifts in the conceptual understanding of material behaviors. A proposed framework delineates the interaction dynamics between surprise signals, algorithmic adaptability, and epistemic feedback loops, emphasizing trade-offs in uncertainty management and knowledge integration. This analysis underscores systems-level insights into how surprise enhances the resilience of discovery pipelines, fostering integrative perspectives on material innovation without positing empirical validations. Ethical considerations arise in interpreting surprise as a catalyst for paradigm evolution, prompting reflections on the epistemic boundaries of AI-assisted science. Overall, this work contributes to a nuanced appreciation of surprise as an intrinsic component in the conceptual architecture of AI-enabled materials research, inviting broader discourse on its interpretive implications.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 7

Scientific Overconfidence in High-Performing Materials AI Systems
The integration of artificial intelligence (AI) into materials science has heightened interpretive challenges regarding model reliability, particularly in systems that exhibit high performance metrics. This conceptual exploration examines the epistemic underpinnings of overconfidence in AI-driven materials predictions, where apparent precision may obscure underlying uncertainties and systemic biases. Drawing from recent literature, the analysis synthesizes how data-driven approaches in materials discovery interact with human cognitive frameworks, fostering interpretive misalignments that influence scientific decision-making. Key dynamics include the interplay between algorithmic robustness and domain-specific knowledge gaps, as well as the feedback structures that perpetuate overreliance on quantitative outputs. Through a proposed framework, the paper interprets these interactions as emergent tensions within socio-technical ecosystems, highlighting ethical considerations in knowledge production. The discussion underscores the need for integrative reasoning that balances technological advancements with epistemic humility, offering insights into steering logics that mitigate distorted interpretations without prescribing empirical validations. Ultimately, this work contributes to a nuanced understanding of how overconfidence manifests in high-stakes AI applications in the materials sciences and advocates reflective practices in scientific inquiry.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 8

Compositional Generalization as a Distinct Failure Mode in Materials AI
The integration of artificial intelligence into materials science has highlighted challenges in model performance, particularly in domains that require extrapolation beyond the training data distribution. This manuscript explores compositional generalization as a unique failure mode in materials AI, in which systems struggle to interpret novel combinations of atomic or molecular elements despite familiarity with individual components. Through a synthesis of recent literature, the analysis delineates how this failure manifests in predictive tasks, such as property estimation in alloys or polymers, revealing underlying tensions between data-driven learning and structural comprehension. Conceptual interpretations highlight the interplay between representational invariance and contextual dependencies, underscoring epistemic gaps in current architectures. The proposed framework interprets these dynamics through lenses of modular interaction and systemic feedback, emphasizing trade-offs in scalability and robustness. By examining the ethical ramifications of deployment in high-stakes applications, the discussion integrates insights into steering mechanisms that could mitigate such limitations without empirical validation. Ultimately, this conceptual inquiry fosters a deeper understanding of AI’s role in advancing materials discovery and advocates for interpretive strategies that prioritize holistic integration over isolated optimizations.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 9

When Models Agree for the Wrong Reasons: A Conceptual Analysis of Consensus in Materials AI
Consensus among machine learning models in materials artificial intelligence often manifests as aligned predictions across ensembles or diverse architectures, yet this alignment frequently conceals underlying misalignments in representational logic or epistemic foundations. This conceptual analysis interprets such phenomena through the lens of interaction dynamics between algorithmic assumptions, uncertainty propagations, and data-systemic interdependencies. By synthesizing insights from recent literature, the discussion illuminates how apparent harmonies in property predictions—such as electronic, mechanical, or thermal attributes—can emerge from shared artifacts rather than a coherent grasp of material phenomena. Analytical implications highlight steering logics in ensemble construction that trade diversity for stability, fostering feedback structures prone to amplifying spurious alignments. Epistemic reasoning underscores the interpretive tension between surface agreement and deeper validation, where consensus serves as an emergent indicator of systemic coherence or fragility. Ethical dimensions arise in the implications for knowledge production in materials discovery, urging nuanced scrutiny to discern integrative fidelity from illusory convergence. The framework advanced here conceptualizes consensus as a multifaceted interpretive construct, shaped by trade-offs in uncertainty handling and model diversity, thereby enriching understanding of AI’s role in reshaping materials’ conceptual landscapes. This approach advocates heightened epistemic vigilance, framing consensus not as a proxy for validation but as a dynamic site for probing the boundaries of interpretive reliability in data-driven materials inquiry.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 10

Latent Variable Leakage in Materials AI: A Conceptual Risk Framework
In the evolving landscape of materials artificial intelligence (AI), latent variables serve as compressed representations that underpin model architectures, facilitating the interpretation of complex material properties and behaviors. This manuscript explores the conceptual dimensions of latent-variable leakage, in which unintended informational flows within these representations may influence systemic outcomes in materials discovery and design. Through an integrative analysis of theoretical underpinnings, the discussion elucidates interaction dynamics between latent spaces and external variables, highlighting epistemic trade-offs in model transparency and generalization. The synthesis of recent literature reveals patterns in how leakage manifests across generative and predictive frameworks, emphasizing steering logics that balance representational fidelity with risk mitigation. A proposed conceptual framework interprets these dynamics as interconnected feedback structures, where leakage pathways intersect with domain-specific constraints in materials science. Ethical reasoning underscores the implications for equitable innovation, while systems-level insights advocate for reflexive approaches in AI deployment. This work contributes to scholarly discourse by framing leakage not as isolated anomalies but as inherent aspects of latent encoding, informing interpretive strategies for sustainable AI integration in materials research.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 11

Failure, Uncertainty, and Risk in Materials AI — How Negative Outcomes Are Handled Across the Literature: A Review Study
The integration of artificial intelligence (AI) and machine learning (ML) in materials science has accelerated discovery and design processes, yet it introduces challenges related to failure, uncertainty, and risk. This narrative review examines how the materials AI literature addresses negative outcomes, including model uncertainties, predictive failures, and associated risks in application. Drawing on peer-reviewed studies, we explore uncertainty quantification techniques, robustness evaluations, and risk mitigation strategies. Key themes include Bayesian methods for uncertainty estimation, benchmark studies on prediction reliability, and strategies to handle data scarcity and extrapolation errors. The review highlights gaps in handling adversarial conditions and real-world failures, proposing future directions for more resilient AI frameworks in materials research. By synthesizing these insights, we aim to foster a more cautious and effective use of AI in advancing materials innovation.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2026 | Article: 95

The Attention Economy of Materials AI: How Model Focus Shapes Scientific Attention Allocation
In the expanding domain of artificial intelligence applied to materials science, computational models do not merely predict properties or accelerate screening; they function as subtle but powerful mechanisms that allocate finite scientific attention across an effectively infinite chemical space. By prioritizing certain compositional regions, structural motifs, or property axes while de-emphasizing others, these systems implicitly decide which questions will be asked, which hypotheses will be tested, and which materials classes will receive downstream experimental or theoretical investment. This position paper argues that Materials AI operates as an attention-allocation infrastructure whose architectural choices reshape the trajectory of discovery itself, transforming what was once an open-ended scientific exploration into a directed economy of focus. Drawing on the well-established “attention economy” metaphor from information systems and cognitive science, we introduce the parallel concept of scientific attention capital—the limited pool of researcher time, funding, instrumentation access, and collective curiosity that models now mediate and, in many cases, ration. Rather than viewing model-induced focus as a neutral technical artifact, we distinguish productive attention (focused investment that yields rapid, high-impact advances in targeted domains) from pathological attention (self-reinforcing loops that create blind spots, reward hacking, and representational injustice). The perspective developed here suggests that recognizing Materials AI as an attention-shaping force carries immediate implications for how the community designs and audits. It deploys these systems if the goal is to preserve the generative openness that has historically driven materials innovation. Ultimately, treating attention allocation as an explicit design variable rather than an incidental byproduct offers a conceptual framework for ensuring that the next generation of Materials AI expands, rather than contracts, the horizons of scientific possibility.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 96

A Conceptual Theory of Model-Science Interface: Where AI Outputs Become Experimental Inputs
In the rapidly evolving field of artificial intelligence for materials science, research has overwhelmingly emphasized the development of predictive models, active learning algorithms, and inverse design strategies to accelerate the identification of novel functional materials. Yet, the critical boundary at which these computational outputs become experimental inputs—the model-science interface—remains largely ignored and treated as an unproblematic transmission step. Existing literature on self-driving laboratories and autonomous experimentation systems, while advancing integrated platforms for clean energy discovery and closed-loop workflows, assumes that model predictions, uncertainty estimates, and experimental recommendations flow seamlessly into synthesis protocols, characterization decisions, and iterative loops without significant distortion or loss. This paper proposes the model-science interface as a distinct object of study, worthy of its own conceptual framework rather than being subsumed under broader discussions of automation or machine learning. By formalizing the interface as the active zone of translation between algorithmic intelligence and empirical practice, the framework distinguishes it from upstream modeling or downstream execution phases, thereby enabling systematic analysis of its internal dynamics. The key concepts articulated herein include a typology of interface operation modes differentiated along dimensions of autonomy and stakes, a detailed examination of information transformations that occur when AI outputs cross into experimental inputs—including preservation of core predictions, loss of contextual nuance, addition of laboratory constraints, and potential distortion through interpretation—and the introduction of “interface fidelity” as a conceptual variable that quantifies the quality of this transition across multiple dimensions. These elements, which build directly upon foundational accounts of autonomous chemical experiments and minimal working examples for self-driving laboratories, provide a vocabulary and set of distinctions for diagnosing interface failure modes that can undermine the overall efficacy of materials discovery pipelines. The framework draws upon foundational ideas in autonomous experimentation while elevating the interface itself as the locus of negotiation between computational promise and physical reality. Ultimately, adopting an interface-aware perspective carries profound implications for materials AI practice. It encourages researchers to design interfaces with intentionality, to report interface specifications alongside model performance, and to study information dynamics explicitly, thereby realizing the full potential of self-driving laboratories for accelerating the discovery of materials for clean energy, piezoelectrics, and beyond. This conceptual contribution thus bridges the persistent gap between model sophistication and experimental impact, fostering more accountable, efficient, and robust autonomous materials research ecosystems.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 97

The Problem of Scientific Regret in AI-Driven Materials Selection
The accelerating integration of artificial intelligence into materials selection processes has brought unprecedented efficiency to high-throughput screening and discovery campaigns, yet it has also introduced a subtle but profound failure mode that remains largely unrecognized in the field: scientific regret. This paper identifies scientific regret as a distinct failure mode in AI-driven materials science—the ex-post realization that a better material or research direction was passed over due to an AI recommendation, often under conditions of irreducible uncertainty and vast combinatorial search spaces. Unlike traditional statistical errors, scientific regret captures the experiential and consequential dimension of missed opportunities in research trajectories that are difficult or impossible to reverse. Drawing on foundational work in decision theory and recent advances in Bayesian optimization for materials discovery, the paper defines scientific regret, delineates its mechanisms of production within AI systems, develops a typology tailored to materials contexts, and outlines principles for its detection and mitigation. By analyzing how premature search space pruning, overconfidence in negative predictions, and misaligned acquisition functions contribute to regret, this analysis reveals how current AI paradigms may systematically undervalue exploration in favor of short-term gains. The implications for materials AI practice are significant, calling for the design of regret-sensitive systems that better balance exploitation with the long-term costs of locked-in choices. Ultimately, embracing scientific regret as a core design constraint promises to foster more robust, reflective, and innovative approaches to autonomous materials research. Scientific regret is not merely an abstract philosophical concern but a practical barrier to genuine progress in materials science. When AI systems guide researchers away from promising chemistries or structures, the subsequent realization of a missed opportunity can stall entire research programs, waste limited experimental resources, and distort the collective knowledge base of the field. This failure mode is especially insidious because materials discovery operates in enormous design spaces where exhaustive enumeration is impossible and where negative predictions are rarely revisited once resources are committed elsewhere. By foregrounding scientific regret as a failure mode, this analysis seeks to reorient the community toward decision frameworks that explicitly account for the irreversible nature of many AI-influenced choices in materials selection.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 98

Conceptual Foundations for Adversarial Validation in Materials Machine Learning
Standard validation protocols in materials machine learning continue to rely on the assumption that training and test data are drawn from the same underlying distribution. This assumption is almost invariably violated in real-world materials datasets because of temporal drift in measurement techniques, compositional biases in database construction, and experimental confounders arising from different laboratories and instruments. This conceptual framework article proposes adversarial validation as a diagnostic tool specifically tailored for materials informatics: a method that trains a discriminator to explicitly detect whether a distribution shift exists between any two datasets, thereby revealing hidden generalization failures that conventional train-test splits and k-fold cross-validation cannot expose. The framework introduces the conceptual foundations of adversarial validation, distinguishes it from adversarial attacks, articulates why the technique is particularly powerful in the small-data, high-dimensional, and physically constrained domain of materials science, and offers a five-component structure for its systematic application—feature-space definition, classifier selection, shift-detection thresholding, localization of driving features, and actionable response rules. By embedding materials-specific domain knowledge into the interpretation of discriminator performance, the approach transforms validation from a passive checkpoint into an active diagnostic that can distinguish temporal shift from compositional bias and experimental confounding. The implications for materials AI practice are immediate and transformative: researchers can now report adversarial validation results alongside standard metrics, trigger targeted dataset augmentation or model retraining when shifts are detected, and document potential sources of distribution mismatch in experimental workflows, ultimately raising the robustness and trustworthiness of property predictions that underpin materials discovery and design.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 99

Algorithmic Forgetting as a Design Choice: A Conceptual Analysis of Memory in Materials AI
The term “forgetting” appears throughout the materials artificial intelligence literature in multiple, often contradictory senses: as a catastrophic failure that destroys previously acquired knowledge of structure–property relations, as an unexamined side effect of data deletion or replay buffer limits, and occasionally as an implicit consequence of model capacity constraints. This conceptual ambiguity impedes precise communication, obscures design decisions, and prevents the field from treating forgetting as a controllable parameter rather than an inevitable defect. The present boundary/definitional paper proposes a precise definition of algorithmic forgetting as a deliberate design choice, distinct from both catastrophic forgetting and passive capacity limits. It distinguishes algorithmic forgetting from five nearby concepts—catastrophic forgetting, data deletion, privacy preservation, capacity saturation, and regularization-induced compression—by clarifying intent, mechanism, epistemic consequences, and reversibility. The paper further articulates the conditions under which forgetting becomes beneficial (adaptation to distribution shift in experimental data streams, selective retention under resource constraints, and controlled deletion for intellectual property or safety) versus harmful (loss of rare but physically valid examples). Finally, it supplies a materials-specific conceptual framework for deciding what to forget and what to retain, grounded in rarity, recency of validation, and relevance to the current search space. By reframing forgetting as an explicit design lever, this analysis offers materials AI practitioners a shared vocabulary and a systematic approach to engineering memory policies that enhance rather than undermine long-term scientific utility.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 100

The Measurement Problem in Materials Informatics: When Observing Changes in the System
In materials informatics, the act of measuring a material property is routinely treated as a neutral act of passive observation. Yet, every measurement consumes finite resources, physically alters the sample, or reshapes the space of future measurements through model-guided selection. This paper identifies a direct analog of the quantum measurement problem within data-driven materials discovery: observation is not merely informative but constitutively changes the system being observed by depleting experimental budgets, inducing material modifications, and biasing the very distribution of data that subsequent AI models will learn. The theoretical claim advanced here is that materials informatics harbors an intrinsic measurement problem in which AI-guided measurement actively constructs rather than neutrally samples the observable landscape, thereby rendering the resulting datasets and models path-dependent on the history of prior observations. Key concepts include resource depletion, selection feedback loops, and measurement-driven evolution, all of which distinguish classical materials measurement effects from quantum collapse while sharing the core epistemic feature of non-neutrality. The implications are far-reaching for AI-guided materials discovery: autonomous laboratories must treat measurement policies as interventions rather than recordings, active-learning algorithms must internalize the cost of altering the observable world, and dataset curation protocols must document measurement history as rigorously as they document final property values. By theorizing this measurement problem, the present analysis offers a conceptual framework that reframes experiment design, model training, and discovery workflows as inherently self-referential processes in which the observer and the observed co-evolve.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 101

A Conceptual Typology of Scientific Surprise in AI-Guided Discovery
The ambiguous usage of the term “surprise” in AI-guided discovery literature represents a significant conceptual barrier in artificial intelligence for materials science. Surprise is variously treated as a statistical anomaly flagged by machine learning models, a human psychological state of unexpectedness that prompts belief revision, an information-theoretic measure of divergence between prior and posterior beliefs, or an unexpected breakthrough that leads to a genuine scientific advance. This lack of precision confuses researchers, fragments the literature, and impedes the systematic design of AI systems capable of deliberately cultivating the forms of unexpectedness that drive materials innovation. This paper proposes a precise typology of scientific surprise consisting of four distinct types—predictive surprise, representational surprise, discovery surprise, and methodological surprise—tailored specifically to the domain of AI-guided discovery in materials science. The key distinctions among these types are articulated along four core dimensions: the source of the surprise (originating in the AI model or in the human scientist), the trigger (prediction error, out-of-distribution data, contradiction with existing theory, or unexpected patterns in the inquiry process itself), the experiencer (primarily the model or the scientist), and the epistemic consequences that follow (model retraining, expansion of representational capacity, theory revision, or redesign of search and measurement strategies). By furnishing this conceptual framework, the paper offers clear implications for designing AI systems that can report, distinguish, and cultivate productive forms of surprise, thereby transforming AI from a passive predictor into an active partner in the discovery process and enabling more effective, targeted responses to different kinds of unexpectedness in materials science.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 102

Default Assumptions as Hidden Parameters: A Theory of Implicit Priors in Materials AI
In the rapidly expanding domain of artificial intelligence applied to materials science, default assumptions embedded within machine learning pipelines—ranging from software library choices and architectural presets to data preprocessing routines and evaluation protocols—are routinely treated as neutral, inconsequential background elements that require no explicit justification. Yet these defaults operate as hidden parameters, subtly yet powerfully constraining the hypothesis space, directing optimization trajectories, and ultimately shaping the predictive behavior of models in ways that rival or even exceed the influence of explicitly tuned parameters, as theoretical analyses of deep networks have long emphasized. This paper advances the theoretical claim that default assumptions in materials AI function as implicit priors, encoding unacknowledged inductive biases that propagate through every stage of a pipeline and determine what counts as a valid or reliable prediction about material properties. Building directly on foundational examinations of inductive bias, we distinguish defaults from both explicit parameters and tunable hyperparameters, develop a taxonomy of four primary default types specific to materials informatics, and derive corollaries concerning the epistemic consequences of unexamined defaults for model comparison, reproducibility, and knowledge transfer. We further examine why such defaults persist—owing to cognitive convenience, historical path dependence, and systematic attribution errors—and clarify their subtle yet critical relation to formal Bayesian priors, while noting that understanding deep learning requires rethinking generalization when defaults remain hidden. The analysis culminates in concrete implications for practice, proposing that defaults must be elevated to first-class objects of documentation, justification, and sensitivity analysis if materials AI is to achieve genuine epistemic transparency and scientific robustness. By theorizing defaults as hidden parameters, this work identifies an overlooked dimension of model epistemology in materials science and offers a conceptual framework for making the invisible visible.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2022 | Article: 103

Scientific Validation in Materials AI—A Critical Survey of Conceptual Approaches: A Review Study
This review systematically surveys conceptual approaches to scientific validation in artificial intelligence applications for materials science, drawing exclusively on 50 peer-reviewed publications from 2017 to 2022 to examine how validation is defined, operationalized, critiqued, and innovated upon within the domain. The methodology followed a targeted literature search protocol across Web of Science, Scopus, and arXiv using eight predefined search strings focused on validation, cross-validation, out-of-distribution testing, generalization, and related terms in materials AI, with strict inclusion criteria requiring explicit discussion of conceptual or epistemological aspects of validation and exclusion of purely empirical performance reports, ultimately yielding the 50 selected references after PRISMA-style screening of approximately 250 unique records. Current validation practices in materials AI literature remain anchored in conventional statistical techniques such as random train-test splits, k-fold cross-validation, leave-one-out cross-validation, and hold-out test sets, which the surveyed papers predominantly employ to quantify predictive accuracy on materials property prediction, discovery, and design tasks. Critical findings demonstrate that these practices frequently claim to establish reliable generalization while actually capturing only in-sample performance, systematically overlooking hidden data structures, distribution shifts, feature selection leakage, and the small-data regimes intrinsic to materials science, thereby producing inflated estimates of model utility that do not translate to real-world deployment. To structure the field’s understanding, the review advances a taxonomy of validation approaches organized hierarchically by what they seek to validate—predictive accuracy, robustness, generalizability, and causal structure—providing a conceptual scaffold for aligning methods with task-specific requirements. Recommendations emphasize explicit reporting, justification of method choice, and community-wide benchmarks. At the same time, open challenges persist in areas such as validating generative models for novelty and enabling trustworthy extrapolation beyond training distributions, underscoring an urgent need for epistemologically grounded practices that match the high-stakes demands of materials discovery.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2022 | Article: 104

The Treatment of Absence and Null Results in Materials Machine Learning Literature: A Review Study
This review systematically examines the treatment of absence and null results in the materials machine learning literature spanning 2017–2022, drawing exclusively on a curated set of 30 peer-reviewed publications and foundational works that address publication bias, negative findings, and reproducibility challenges in data-driven materials discovery. Through a targeted search strategy across databases such as Web of Science, Scopus, and arXiv using terms including “null result,” “negative result,” “publication bias,” “file drawer,” “failed synthesis,” and “reproducibility” combined with materials informatics keywords, the analysis reveals a persistent imbalance: while successful predictions and syntheses dominate published outputs, systematic documentation of failed predictions, unsuccessful syntheses, null correlations, and abandoned model architectures remains exceedingly rare. What is currently reported tends to be limited to negative outcomes that coincidentally reveal mechanistic insights or contradict high-profile hypotheses, whereas what is systematically unreported encompasses the vast majority of unsuccessful hyperparameter searches, negative active learning campaigns, and non-discoveries that yield no novel materials meeting target criteria. The typology of absence and null results developed here identifies six distinct categories—negative predictive outcomes, null hypothesis non-rejection, failed synthesis, non-discovery, failed replication, and abandoned architecture—each carrying unique implications for scientific progress. The consequences of this non-reporting include severe overestimation of model performance, widespread redundant experimental effort, a false sense of methodological consensus across the field, and slowed overall discovery rates as potentially informative negative signals remain invisible. Ultimately, this review offers concrete recommendations for authors, journals, and the broader community to shift incentives toward transparent reporting of absence, thereby restoring balance to the materials AI literature and accelerating reliable data-driven discovery.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2022 | Article: 105

The Rise of Platform-Based Competition: A Review of Theoretical Perspectives on Digital Marketplaces, Network Effects, and Ecosystem Strategy
Platform-based competition has fundamentally altered the nature of rivalry in digital markets, shifting emphasis from firm-level resources to network effects, multi-sided participation, and ecosystem orchestration. This integrative review synthesizes theoretical perspectives on digital marketplaces, network effects, and ecosystem strategy, drawing on 35 peer-reviewed sources published between 2003 and 2026. It examines how platform market structures differ from traditional competition, the mechanisms through which network effects generate scaling advantages and competitive lock-in, and the strategic role of governance in balancing openness with control. The analysis highlights complementor dynamics, value creation versus capture tensions, and the evolving interplay between platform leaders, users, and complementors. By classifying and comparing core theoretical streams, the review identifies persistent strategic tensions—openness versus control, scale versus governance complexity, and innovation versus appropriation—and traces the maturation of the field from early two-sided market models to contemporary ecosystem perspectives. To advance coherence, the review introduces the platform competition layered synthesis (PCLS) model, a novel integrative architecture that organizes the literature into six interconnected layers. The model reveals feedback mechanisms through which market outcomes continuously reshape platform design and competitive positioning. Implications for digital business strategy and future research directions are discussed.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 July 2022 | Article: 106

Discovery without Understanding: A Systems Theory of Black-Box Optimization in Autonomous Materials Engineering
In the evolving landscape of computational and data-driven materials engineering, the integration of machine learning and high-throughput methodologies has accelerated discovery processes, yet it introduces a paradox where rapid optimization often bypasses deep scientific understanding. This manuscript presents a systems theory perspective on black-box optimization in autonomous materials engineering, emphasizing closed-loop labs where AI-driven decisions guide experimentation without explicit interpretability. Drawing from materials informatics and representation learning, we identify the discovery acceleration paradox: enhanced efficiency in inverse design and property prediction erodes traditional epistemic structures, leading to reliance on opaque models. We introduce the "Epistemic Opaque Discovery System" (EODS) framework, which conceptualizes materials discovery as a layered network of data infrastructures, model architectures, and feedback mechanisms. This framework highlights trade-offs between optimization speed and interpretability, incorporating uncertainty quantification to mitigate risks in autonomous systems. Implications extend to simulation-experiment coupling and multimodal datasets, suggesting pathways for balanced computational workflows that preserve scientific insight amid black-box dominance. By reframing discovery pipelines, EODS offers a theoretical lens for engineering resilient AI ecosystems in materials science, fostering sustainable innovation without sacrificing foundational knowledge.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 76

Representation Is Not Reality: Epistemic Limits of Learned Materials Embeddings in Computational Design Systems
The rapid evolution of computational and data-driven materials engineering has transformed materials discovery from traditional trial-and-error approaches to sophisticated AI-integrated pipelines. Within this paradigm, learned embeddings serve as foundational representations that encode complex material properties, structures, and behaviors into latent spaces amenable to machine learning algorithms. However, these embeddings, while powerful for predictive modeling and high-throughput screening, introduce epistemic limits that challenge the fidelity of computational design systems. This manuscript explores the disconnect between representational abstractions and physical reality, emphasizing how embedding-induced biases, dimensionality reductions, and generalization assumptions constrain the reliability of AI-guided materials innovation. We introduce a novel conceptual framework, the Epistemic Representation Cascade (ERC), which dissects the multi-layered interactions between data infrastructures, learning architectures, and discovery workflows to reveal inherent epistemic risks. By integrating insights from materials informatics and representation learning, the ERC highlights feedback mechanisms that amplify or mitigate these limits, offering systems-level guidance for enhancing interpretability and robustness in autonomous design ecosystems. Implications extend to closed-loop experimentation and inverse design, advocating for infrastructure-aware strategies that prioritize epistemic alignment over mere predictive accuracy. This work underscores the need for balanced computational steering in materials AI, fostering more trustworthy pathways for next-generation materials engineering.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 77

Scaling Laws without Physics: A Conceptual Analysis of Model Expansion in Computational Materials Engineering
The rapid evolution of computational materials engineering has ushered in an era where data-driven approaches increasingly dominate discovery pipelines, leveraging vast datasets and expansive model architectures to uncover material properties and behaviors. This conceptual analysis examines the phenomenon of model expansion in materials informatics, focusing on scaling laws that emerge independently of traditional physics-based derivations. By dissecting the interplay between dataset scaling, parameter proliferation, and computational resource demands, we highlight how such expansions influence epistemic gains in materials discovery. A core gap in current paradigms lies in the overreliance on empirical scaling metrics, which often overlook the nuanced trade-offs between model complexity and interpretive insight. To address this, we introduce the "Insight Amplification Cascade" framework, a layered conceptual structure that maps data infrastructures to inference dynamics, emphasizing feedback mechanisms that balance energy costs against discovery yields. This framework integrates representation learning with uncertainty quantification to steer computational workflows toward sustainable scaling. Implications extend to autonomous discovery systems, where model expansion fosters robust inverse design without necessitating physics-grounded priors. Ultimately, this analysis underscores the need for infrastructure-level reforms in materials AI, promoting scalable yet interpretable ecosystems that enhance long-term innovation in computational materials engineering. Through this lens, we advocate for a reevaluation of scaling strategies to prioritize epistemic efficiency over mere parametric growth.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 78

When Data Steers Design: Feedback Dynamics in AI-Guided Materials Exploration Pipelines
The integration of computational tools and data-driven methodologies has transformed materials engineering, enabling accelerated discovery through AI-assisted pipelines that link data acquisition, model training, and experimental validation. In this paradigm, materials informatics leverages vast datasets from high-throughput computations and multimodal sources to inform design decisions, yet inherent feedback dynamics often introduce biases that steer exploration trajectories in unintended ways. This conceptual manuscript identifies a critical gap in understanding how data-model-experiment loops can self-reinforce certain pathways, leading to narrowed exploration spaces and amplified discovery biases. To address this, we introduce the Feedback Steering Framework (FSF), a systems-level architecture that interprets the interplay between data representations, model inferences, and iterative design cycles. The framework elucidates mechanisms such as reinforcement discovery bias, where initial data patterns perpetuate model preferences, and exploration narrowing, wherein computational steering logics constrain the search space over successive iterations. By conceptualizing these dynamics, FSF provides insights into optimizing AI-guided materials exploration for broader epistemic coverage. Implications extend to computational materials science ecosystems, including enhanced uncertainty management in autonomous systems and more robust inverse design strategies, ultimately fostering resilient infrastructures for next-generation materials innovation. This work underscores the need for interpretive tools that balance computational efficiency with comprehensive discovery potential in data-steered environments.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 79

Material Spaces Are Not Euclidean: A Computational Critique of Distance Metrics in Data-Driven Materials Discovery
In the rapidly evolving field of computational materials engineering, data-driven approaches have transformed the discovery and design of novel materials by leveraging machine learning and high-throughput computations to navigate vast chemical spaces. Traditional methodologies often rely on Euclidean distance metrics to quantify similarities between materials in latent representations, facilitating tasks such as property prediction, inverse design, and autonomous experimentation. However, this assumption overlooks the inherent non-linearities and topological complexities of material spaces, where properties like electronic bandgaps, mechanical strengths, and thermodynamic stabilities emerge from intricate atomic interactions that do not conform to flat geometries. This conceptual gap leads to inefficiencies in representation learning, biased uncertainty quantification, and suboptimal steering in discovery pipelines. Here, we introduce a novel interpretive framework that critiques Euclidean metrics through a manifold-based lens, emphasizing geodesic distances and curvature-aware embeddings to better capture the epistemic structure of materials data. By integrating insights from graph neural networks, multimodal datasets, and closed-loop systems, this framework reveals computational trade-offs in data infrastructures and enhances the interpretability of AI-guided workflows. Implications extend to improved coupling of simulations and experiments, fostering more robust foundation models for materials science and accelerating innovation in energy, electronics, and structural applications without empirical validation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 80

From High-Throughput Computation to Autonomous Discovery: A Review of Closed-Loop Data Infrastructures in Materials Engineering
The field of materials engineering has undergone a profound transformation through the integration of high-throughput computation and data-driven methodologies, evolving from traditional trial-and-error approaches to sophisticated closed-loop systems that accelerate discovery. This review synthesizes recent advancements in computational and data-driven materials ecosystems, focusing on the infrastructure enabling autonomous discovery. Key elements include materials informatics platforms that leverage machine learning for property prediction and inverse design, graph neural networks for representation learning, and high-throughput computational workflows that generate multimodal datasets. We examine the progression from static high-throughput screening to dynamic, closed-loop paradigms incorporating active learning, uncertainty quantification, and simulation-experiment integration. Autonomous laboratories represent a pinnacle of this evolution, where AI orchestrates iterative cycles of hypothesis generation, experimentation, and refinement. The synthesis highlights how these infrastructures bridge computational predictions with experimental validation, fostering inverse materials design and optimizing resource allocation in complex chemical spaces. Challenges in data interoperability and model generalizability are noted, alongside prospects for scalable, self-optimizing systems. Overall, this review positions closed-loop data infrastructures as foundational to next-generation materials engineering, promising accelerated innovation in areas like energy storage, catalysis, and structural materials. By integrating diverse literature, we provide a systems-level perspective on how these tools are reshaping the discovery landscape.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 March 2022 | Article: 81

Representation Learning in Materials Science: Architectures, Data Modalities, and Discovery Applications
The field of materials science has undergone a transformative shift with the integration of computational and data-driven approaches, particularly through representation learning techniques that enable efficient handling of complex materials data. This review synthesizes recent advancements in architectures for representation learning, encompassing graph neural networks, attention-based models, and physics-inspired embeddings, which facilitate the extraction of meaningful features from diverse data modalities such as atomic structures, stoichiometries, and spectroscopic data. By bridging traditional computational methods with machine learning, these representations have accelerated property prediction, inverse design, and materials discovery applications, addressing challenges in high-dimensional spaces and sparse datasets. The scope of this narrative review covers the evolution from basic informatics to sophisticated multimodal integrations, highlighting how data ecosystems and learning frameworks contribute to autonomous discovery pipelines. A systems-level perspective is adopted to integrate cross-study insights, revealing synergies between representation learning and closed-loop systems that couple simulations with experiments. Looking ahead, the review posits that continued refinement of these architectures will drive scalable, AI-guided materials engineering, fostering innovations in energy, electronics, and structural materials while emphasizing the need for robust, interpretable models in real-world applications.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 March 2022 | Article: 82
Filters
Clear All





Access type