Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Algorithmic Confidence as a Control Signal in Materials Research
Materials research increasingly relies on machine learning to accelerate property prediction and discovery, yet the trustworthiness of these models remains constrained by their inability to express epistemic limitations. Algorithmic confidence—embodied in principled uncertainty quantification—provides a quantitative measure of model reliability that can extend beyond diagnostic assessment to serve as an active control signal within the research process. This conceptual manuscript synthesizes recent developments in uncertainty-aware machine learning, Bayesian approaches, and adaptive sampling strategies to argue that confidence estimates hold untapped potential as dynamic regulators of investigative workflows. Rather than treating uncertainty solely as a performance metric or sampling criterion, we conceptualize it as a central control variable that modulates decision pathways, balances exploration and exploitation, and informs the transition from computational prediction to empirical validation. A novel framework is proposed wherein algorithmic confidence governs iterative cycles in materials inquiry, enabling self-regulating mechanisms that align model assertions with epistemic boundaries. This perspective reframes uncertainty not as a limitation but as a strategic operator capable of guiding resource-efficient, robust materials exploration in a purely conceptual sense. By elevating confidence to a control role, the approach seeks to foster more deliberate and principled integration of computational intelligence into materials science paradigms.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2023 | Article: 21

The Cost of Early Convergence in AI-Guided Materials Search
AI-guided materials search increasingly relies on probabilistic and adaptive algorithms to navigate complex design spaces. Within these processes, early convergence emerges as a recurring dynamic wherein search trajectories stabilize around promising regions before exhaustive mapping of uncertainty landscapes occurs. This conceptual manuscript examines the interpretive costs associated with such premature stabilization, framing them not as isolated computational inefficiencies but as interconnected epistemic, structural, and innovation-limiting phenomena. Drawing on recent developments in Bayesian optimization, active learning, and equivariant graph representations for materials systems, the analysis examines how early convergence interacts with exploration-exploitation trade-offs, cascading into effects on knowledge breadth and discovery potential. A novel conceptual framework is advanced that conceptualizes these dynamics through feedback loops and trade-off structures, emphasizing systems-level insights into how algorithmic steering logics shape long-term trajectories in materials innovation. By focusing exclusively on interpretive and integrative dimensions, the contribution highlights the need for refined conceptual models that account for hidden costs embedded in convergence behaviors. This perspective encourages deeper reflection on the epistemic foundations of AI-assisted discovery without invoking empirical validation or predictive assertions.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2023 | Article: 24

Active Learning-Driven Bayesian Optimization of Catalytic Nanoparticles for CO₂ Reduction
The escalating global challenge of carbon dioxide (CO₂) emissions necessitates innovative approaches to mitigate climate change through efficient catalytic conversion. This conceptual manuscript proposes a novel theoretical framework that integrates active learning with Bayesian optimization to enhance the design of catalytic nanoparticles for CO₂ reduction. Drawing on principles from machine learning and materials science, the framework addresses the complexities of high-dimensional parameter spaces in nanoparticle synthesis, such as size, shape, composition, and surface facets, which influence catalytic performance. By leveraging active learning to intelligently select informative data points and Bayesian optimization to refine surrogate models iteratively, the approach theoretically accelerates the identification of optimal nanoparticle configurations without empirical validation. The framework emphasizes uncertainty quantification and adaptive sampling to efficiently navigate the vast design space. This synthesis of concepts from recent literature highlights gaps in traditional optimization methods and posits that the proposed integration could conceptually reduce exploration costs while enhancing selectivity and activity in CO₂ reduction processes. The manuscript outlines theoretical underpinnings, a proposed framework, and implications for applied artificial intelligence in materials science, fostering future conceptual advancements in sustainable catalysis.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2023 | Article: 34

Recent Advances in Machine Learning-Accelerated Materials Discovery — From Descriptors to Autonomous Experiments
Machine learning (ML) has become a central driver of modern materials discovery, fundamentally reshaping how materials are designed, screened, and experimentally realized. This review examines recent advances in ML-accelerated materials discovery and emphasizes the ongoing progress in material representation and descriptor development toward fully autonomous experimental platforms. We discuss how increasingly sophisticated descriptors—ranging from composition-based features and structure-aware representations to ab initio–derived and learned embeddings—have improved predictive accuracy, data efficiency, and physical interpretability across diverse materials systems. Based on these findings, we discuss the evolution of ML frameworks for property prediction, classification, and inverse design, with particular attention to uncertainty-aware modeling, multiobjective optimization, and explainable learning strategies that bridge predictive performance with scientific insight. The study also highlights the growing role of active learning and generative models in efficiently navigating vast chemical and structural spaces, enabling data-efficient exploration and hypothesis-driven discovery. At the frontier of these developments, autonomous experimental systems integrate ML with robotics to form closed-loop workflows that iteratively design, execute, and refine experiments with minimal human intervention. Applications spanning perovskites, alloys, energy materials, and nanostructures illustrate the broad impact of these approaches in overcoming traditional trial-and-error limitations. Finally, we discuss persistent challenges associated with data scarcity, extrapolation, interpretability, and system integration, and outline future directions toward more robust, scalable, and sustainable autonomous materials discovery. Collectively, these advances represent a paradigm shift from passive data-driven prediction to intelligent, self-guided materials innovation.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2024 | Article: 41

Uncertainty-Conditioned Experiment Planning: A Conceptual Framework for AI-Guided Materials Exploration
Materials exploration faces persistent challenges stemming from vast chemical spaces, high experimental costs, and inherent uncertainties in predictive models. While machine learning has accelerated property prediction and guided candidate selection, conventional approaches often treat uncertainty as a uniform metric within fixed acquisition strategies. This conceptual paper introduces uncertainty-conditioned experiment planning (UCEP) as a novel theoretical framework for AI-guided materials discovery. UCEP reframes experiment planning as a dynamic process conditioned on the multidimensional character of uncertainty, integrating epistemic and aleatoric components, data-related biases, and model limitations into the steering logic. Rather than relying on static acquisition functions, the framework emphasizes adaptive interaction dynamics between uncertainty characterization and planning decisions, enabling context-sensitive trade-offs between exploration, exploitation, and bias mitigation. Drawing on interpretive insights from materials informatics and uncertainty quantification literature, UCEP highlights systems-level feedback structures that can enhance epistemic robustness and scientific efficiency without presupposing empirical outcomes. The framework offers analytical implications for rethinking how AI systems interpret and respond to uncertainty in iterative discovery cycles, contributing to more reflective and integrative AI-assisted materials research.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 88

Small-Data and Sparse-Regime Learning in Materials AI — Methods, Assumptions, and Limits
The integration of artificial intelligence (AI) and machine learning (ML) into materials science, often referred to as materials informatics or materials AI, has accelerated the discovery, design, and optimization of advanced materials. However, materials science frequently operates in small-data and sparse-regime conditions, where datasets are limited in size (often tens to hundreds of samples), high-dimensional, imbalanced, or sparsely populated due to the high cost, time, and complexity of experimental measurements and high-fidelity simulations. This narrative review synthesizes recent advances in methods tailored to these constraints, categorizing approaches at the data-source level (e.g., literature extraction, database construction, high-throughput workflows), algorithmic level (e.g., support vector machines, Gaussian process regression, ensemble models, imbalanced learning techniques), and strategic level (e.g., active learning, transfer learning). Key assumptions underlying these methods are examined, including similarity between source and target domains for transfer learning, representativeness of initial samples and reliable uncertainty quantification in active learning, and the validity of physical priors or inductive biases in physics-informed approaches. The review also addresses inherent limits, such as risks of overfitting, poor generalization beyond the training distribution, sensitivity to data quality and noise, challenges in uncertainty calibration, and dependence on domain expertise. By highlighting successful applications in property prediction, alloy design, and perovskite optimization, this work elucidates the current capabilities and boundaries of small-data and sparse-regime learning in materials AI, guiding researchers navigating data-limited environments.
Journal of Artificial Intelligence for Materials Science
Review | Open access | 18 January 2026 | Article: 90

The Attention Economy of Materials AI: How Model Focus Shapes Scientific Attention Allocation
In the expanding domain of artificial intelligence applied to materials science, computational models do not merely predict properties or accelerate screening; they function as subtle but powerful mechanisms that allocate finite scientific attention across an effectively infinite chemical space. By prioritizing certain compositional regions, structural motifs, or property axes while de-emphasizing others, these systems implicitly decide which questions will be asked, which hypotheses will be tested, and which materials classes will receive downstream experimental or theoretical investment. This position paper argues that Materials AI operates as an attention-allocation infrastructure whose architectural choices reshape the trajectory of discovery itself, transforming what was once an open-ended scientific exploration into a directed economy of focus. Drawing on the well-established “attention economy” metaphor from information systems and cognitive science, we introduce the parallel concept of scientific attention capital—the limited pool of researcher time, funding, instrumentation access, and collective curiosity that models now mediate and, in many cases, ration. Rather than viewing model-induced focus as a neutral technical artifact, we distinguish productive attention (focused investment that yields rapid, high-impact advances in targeted domains) from pathological attention (self-reinforcing loops that create blind spots, reward hacking, and representational injustice). The perspective developed here suggests that recognizing Materials AI as an attention-shaping force carries immediate implications for how the community designs and audits. It deploys these systems if the goal is to preserve the generative openness that has historically driven materials innovation. Ultimately, treating attention allocation as an explicit design variable rather than an incidental byproduct offers a conceptual framework for ensuring that the next generation of Materials AI expands, rather than contracts, the horizons of scientific possibility.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 96

A Conceptual Theory of Model-Science Interface: Where AI Outputs Become Experimental Inputs
In the rapidly evolving field of artificial intelligence for materials science, research has overwhelmingly emphasized the development of predictive models, active learning algorithms, and inverse design strategies to accelerate the identification of novel functional materials. Yet, the critical boundary at which these computational outputs become experimental inputs—the model-science interface—remains largely ignored and treated as an unproblematic transmission step. Existing literature on self-driving laboratories and autonomous experimentation systems, while advancing integrated platforms for clean energy discovery and closed-loop workflows, assumes that model predictions, uncertainty estimates, and experimental recommendations flow seamlessly into synthesis protocols, characterization decisions, and iterative loops without significant distortion or loss. This paper proposes the model-science interface as a distinct object of study, worthy of its own conceptual framework rather than being subsumed under broader discussions of automation or machine learning. By formalizing the interface as the active zone of translation between algorithmic intelligence and empirical practice, the framework distinguishes it from upstream modeling or downstream execution phases, thereby enabling systematic analysis of its internal dynamics. The key concepts articulated herein include a typology of interface operation modes differentiated along dimensions of autonomy and stakes, a detailed examination of information transformations that occur when AI outputs cross into experimental inputs—including preservation of core predictions, loss of contextual nuance, addition of laboratory constraints, and potential distortion through interpretation—and the introduction of “interface fidelity” as a conceptual variable that quantifies the quality of this transition across multiple dimensions. These elements, which build directly upon foundational accounts of autonomous chemical experiments and minimal working examples for self-driving laboratories, provide a vocabulary and set of distinctions for diagnosing interface failure modes that can undermine the overall efficacy of materials discovery pipelines. The framework draws upon foundational ideas in autonomous experimentation while elevating the interface itself as the locus of negotiation between computational promise and physical reality. Ultimately, adopting an interface-aware perspective carries profound implications for materials AI practice. It encourages researchers to design interfaces with intentionality, to report interface specifications alongside model performance, and to study information dynamics explicitly, thereby realizing the full potential of self-driving laboratories for accelerating the discovery of materials for clean energy, piezoelectrics, and beyond. This conceptual contribution thus bridges the persistent gap between model sophistication and experimental impact, fostering more accountable, efficient, and robust autonomous materials research ecosystems.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 97

The Measurement Problem in Materials Informatics: When Observing Changes in the System
In materials informatics, the act of measuring a material property is routinely treated as a neutral act of passive observation. Yet, every measurement consumes finite resources, physically alters the sample, or reshapes the space of future measurements through model-guided selection. This paper identifies a direct analog of the quantum measurement problem within data-driven materials discovery: observation is not merely informative but constitutively changes the system being observed by depleting experimental budgets, inducing material modifications, and biasing the very distribution of data that subsequent AI models will learn. The theoretical claim advanced here is that materials informatics harbors an intrinsic measurement problem in which AI-guided measurement actively constructs rather than neutrally samples the observable landscape, thereby rendering the resulting datasets and models path-dependent on the history of prior observations. Key concepts include resource depletion, selection feedback loops, and measurement-driven evolution, all of which distinguish classical materials measurement effects from quantum collapse while sharing the core epistemic feature of non-neutrality. The implications are far-reaching for AI-guided materials discovery: autonomous laboratories must treat measurement policies as interventions rather than recordings, active-learning algorithms must internalize the cost of altering the observable world, and dataset curation protocols must document measurement history as rigorously as they document final property values. By theorizing this measurement problem, the present analysis offers a conceptual framework that reframes experiment design, model training, and discovery workflows as inherently self-referential processes in which the observer and the observed co-evolve.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 101

Algorithmic Settling Time: A Conceptual Framework for When Materials AI Outputs Stabilize
In the rapidly expanding domain of Artificial Intelligence for Materials Science, researchers routinely train machine learning models until training loss appears to converge. Yet, this practice overlooks a critical and distinct phenomenon: the point at which model outputs themselves cease to change meaningfully with further iterations or data. Algorithmic settling time is introduced here as the number of training iterations, epochs, data points, or active-learning cycles after which predictions for a given input distribution stabilize within a predefined tolerance, independent of loss minimization. This conceptual framework highlights five key factors—data scarcity, feature dimensionality, model complexity, task difficulty, and optimization dynamics—that modulate settling behavior in materials contexts where datasets are sparse, and property landscapes are high-dimensional. A four-component framework for settling-time analysis is proposed, centered on output monitoring, tolerance specification, settling detection, and confidence assessment, offering a principled alternative to ad-hoc early stopping. By foregrounding settling time as an overlooked parameter, this framework promises to enhance reproducibility, reduce computational waste, and improve the reliability of materials predictions ranging from crystal-property regression to generative molecular design, ultimately elevating the epistemic rigor of Materials AI practice.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2023 | Article: 109

A Conceptual Distinction between Exploration Noise and Scientific Error
In the field of artificial intelligence applied to materials science, a fundamental conflation persists in which exploration noise and scientific error are routinely conflated as interchangeable “mistakes” that must be minimized or eliminated to improve model performance. This paper proposes precise conceptual definitions that separate exploration noise—understood as stochastic variation deliberately or unavoidably introduced into decision-making processes to probe uncertain regions of materials design space—from scientific error, defined as any deviation from ground truth that reduces predictive fidelity, distorts mechanistic understanding, or precipitates incorrect materials decisions without any compensating epistemic gain. The distinction matters profoundly because the systematic elimination of exploration noise eradicates the very mechanism that drives discovery in high-dimensional, data-scarce materials landscapes. In contrast, misclassifying scientific error as mere noise allows systematic flaws to propagate undetected through autonomous discovery pipelines. To resolve this ambiguity, the present work offers a four-criterion framework grounded in intentionality, epistemic benefit, systematicity, and correctability that enables researchers to classify any observed deviation with conceptual clarity. Adoption of this framework carries immediate implications for materials AI practice: it demands new reporting standards that explicitly quantify and justify exploration noise, revised peer-review criteria that interrogate rather than penalize productive randomness, and a cultural shift that reframes stochasticity not as a defect to be denoised but as an essential epistemic resource for accelerating the discovery of novel materials with targeted functionalities.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 149

Separating What We Know from What We Guess: A Modular Framework for Epistemic vs. Aleatoric Uncertainty in ML Potentials
Uncertainty quantification in machine learning (ML) interatomic potentials remains fundamentally limited by the conflation of epistemic uncertainty, arising from incomplete sampling of configuration space, and aleatoric uncertainty, embedded in reference data generated by density-functional theory. Existing approaches provide internally consistent uncertainty estimates but collapse these distinct sources into a single scalar, obscuring the mechanisms governing model reliability and limiting principled decision-making. This work introduces a modular, architecture-agnostic framework that enforces explicit separation of epistemic and aleatoric contributions at the level of model design rather than post hoc analysis. The framework defines five interoperable components—a shared feature extractor, dedicated epistemic and aleatoric modules, an aggregation mechanism, and a calibration stage—whose interactions preserve disentanglement throughout training, inference, and downstream application. The resulting formulation transforms uncertainty into an operational diagnostic. Epistemic uncertainty identifies regions where additional data acquisition is informative, whereas aleatoric uncertainty defines the intrinsic accuracy ceiling imposed by the reference method. This separation restructures active learning by directing sampling toward reducible error, enables meaningful comparison between models through their uncertainty composition, and grounds performance evaluation relative to an explicit noise floor. The framework further introduces operational criteria that provide falsifiable tests of successful separation, ensuring that reported uncertainties remain interpretable and consistent across architectures. By decoupling learnable structure from irreducible variability, the proposed approach establishes a principled foundation for uncertainty-aware ML potentials, supporting more efficient data allocation, more reliable atomistic simulations, and more rigorous standards for model development in computational materials science.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2023 | Article: 15

The Interpolation–Extrapolation Boundary in ML Potentials for Phase Space Sampling: A Definitional Framework
The reliability of machine-learned (ML) interatomic potentials in phase space sampling depends critically on whether sampled configurations lie within the interpolative domain of the training data or extend into extrapolative regimes. Despite its central importance, the interpolation–extrapolation distinction remains inconsistently defined across the literature. Existing single-metric approaches—such as convex-hull composition checks, descriptor-space distances, uncertainty estimates, and local-environment similarity—capture only partial aspects of high-dimensional configuration space and frequently yield conflicting classifications. This lack of a rigorous, unified boundary undermines active learning strategies, uncertainty quantification, benchmarking, and the safe deployment of ML potentials in safety-critical applications. This boundary/definitional article introduces a unified and operational framework that delineates the interpolation–extrapolation boundary through four independent and jointly necessary criteria: (C1) compositional coverage, (C2) local-environment similarity, (C3) configurational co-occurrence, and (C4) thermodynamic condition. A configuration is defined as interpolative only when all four criteria are simultaneously satisfied; violation of any criterion constitutes extrapolation. To replace binary classification, the framework further introduces a graded hierarchy of extrapolation—mild, moderate, and severe—based on the number and magnitude of violations. The proposed definitions are deliberately operational, relying exclusively on information available during training-set construction or simulation runtime, and are agnostic to model architecture. Their adoption enables standardized reporting, stratified benchmarking, improved calibration of uncertainty estimates, and more targeted active-learning workflows. By establishing a clear and reproducible boundary, the framework provides a principled foundation for evaluating model reliability during phase space exploration. This work advances the epistemic rigor of ML-driven materials modeling and supports the responsible and transparent deployment of data-driven interatomic potentials.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2023 | Article: 24

Why Active Learning Fails for Defect Formation Energies: When Uncertainty Sampling Misleads Discovery
Active learning is widely used in materials discovery to reduce the number of expensive density functional theory calculations required for surrogate-model training. Its dominant acquisition rule, uncertainty sampling, assumes that the most uncertain configuration will also yield the greatest information gain. Although this logic performs well for many bulk properties, it breaks down for defect formation energies. In this setting, uncertainty sampling repeatedly selects uninformative structures, misallocates computational budget, and fails to reach the defective configurations that govern vacancy, interstitial, and substitutional behavior in technologically important materials. This failure reflects a structural mismatch between the acquisition rule and defect physics. Defect configurations are rare in configuration space, energetics are highly localized, and DFT labels often contain substantial aleatoric noise from finite-size effects, charge corrections, and supercell artifacts. These conditions weaken the link between predictive uncertainty and useful learning signal, while representation bias in models trained mainly on perfect crystals further erodes selectivity. This study develops a failure-mode analysis of uncertainty sampling for defect formation energies, identifying recurrent breakdowns in sampling, representation, calibration, and budget use. It also outlines practical detection principles and defect-aware mitigation strategies, including initialization with defective structures, hybrid acquisition functions, epistemic-only scoring, and budget partitioning. The central implication is that active learning for defect discovery cannot rely on generic uncertainty-based querying. Effective acceleration in this domain requires acquisition strategies designed around the rarity, locality, and noise structure of defect energetics.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2024 | Article: 30

Self-Consistency Is Not Enough: A Framework for Error-Correcting ML Potentials with On-the-Fly Residual Learning
Machine learning interatomic potentials now enable molecular dynamics simulations with near-density-functional-theory accuracy at scales inaccessible to conventional quantum methods. Yet most remain fundamentally static: once trained, they are deployed without adaptation, even as simulations enter configurations outside the original training distribution. Although self-consistency ensures that predicted forces remain exact derivatives of the learned energy surface, it guarantees only internal coherence, not fidelity to the true reference landscape. As a result, small local errors can accumulate into substantial long-timescale drift. This paper proposes a conceptual framework for error-correcting machine learning potentials based on on-the-fly residual learning. The architecture combines a self-consistent base predictor with an error detector, a lightweight residual corrector, an online updater, and a memory manager. Embedded directly within the molecular dynamics loop, these components enable the system to identify unreliable predictions, apply immediate corrections, selectively request sparse density-functional-theory labels, and retain corrective knowledge during continuous adaptation. By shifting from static deployment to simulation-aware error correction, the framework addresses the central limitations of extrapolation failure and accumulated drift. It therefore outlines a path toward adaptive machine learning potentials capable of sustaining reliable long-timescale materials simulations with controlled computational overhead.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2024 | Article: 36

What Does “Sparsity” Mean in High-Dimensional Materials Design? A Boundary Problem for Active Learning
The term “sparsity” is widely used in high-dimensional materials design, yet its meaning remains conceptually unstable and methodologically consequential. This article argues that sparsity in materials discovery should not be treated as a single condition, because the challenges faced by active-learning systems arise from qualitatively different sources. A boundary framework is introduced that distinguishes data sparsity, representation sparsity, and coverage sparsity, and relates each to a distinct failure mode in surrogate modeling and acquisition design. Data sparsity describes the imbalance between the number of evaluated materials and the effective volume of the design space; representation sparsity concerns zero-dominated descriptors; coverage sparsity captures the uneven spatial distribution of samples across composition or descriptor space. By separating these regimes, the analysis shows why standard acquisition functions often underperform in realistic campaigns: they assume sufficient support for interpolation, manageable descriptor structure, and reasonably uniform sampling. The article formulates operational criteria for diagnosing each sparsity type and demonstrates how strategy selection should change as the dominant constraint shifts. In doing so, it provides a clearer vocabulary for high-dimensional materials design and a practical basis for more reliable, sparsity-aware active-learning workflows.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2024 | Article: 38

Flawed Assumptions in Uncertainty Quantification for High-Entropy Alloys: A Critique of Ensemble Diversity Metrics
Ensemble methods have become the dominant framework for uncertainty quantification in machine learning models for high-entropy alloy (HEA) property prediction, where variance across independently trained neural networks is routinely interpreted as epistemic uncertainty. This metric now underpins active learning, compositional screening, and experimental decision-making, largely due to its simplicity and success in data-rich domains. This work shows that such reliance is fundamentally misplaced in HEAs. Ensemble variance implicitly assumes IID sampling, feature-space isotropy, uniform error, independence among models, and Gaussian residuals—conditions that are systematically violated in compositionally complex alloys. HEA datasets are biased toward equiatomic, stable compositions, the compositional manifold is anisotropic, predictive error is strongly heteroscedastic, ensemble members exhibit correlated failures, and extrapolation induces heavy-tailed errors. Under these conditions, ensemble variance becomes miscalibrated, underestimating uncertainty in sparse regions while overstating model reliability. The resulting distortions propagate through discovery workflows, yielding inefficient active-learning strategies, overconfident extrapolation, and misleading experimental guidance, even as larger ensembles appear to reduce uncertainty without improving accuracy. By linking these failures to their statistical origins, this paper clarifies why standard ensemble diversity is not a faithful proxy for epistemic uncertainty in HEAs. It further outlines diagnostics and targeted corrections that align uncertainty estimates with the physics and data structure of complex alloy systems, enabling more reliable and efficient materials discovery.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2024 | Article: 40

Autonomous Laboratories and Closed-Loop Materials Discovery: A Conceptual Framework for Self-Driving Experimentation
Materials discovery remains one of the central bottlenecks in the translation of scientific ideas into functional technologies. Conventional workflows often depend on slow cycles of hypothesis, synthesis, testing, interpretation, and redesign. These cycles are constrained by human time, manual dexterity, instrument availability, and the practical impossibility of exploring vast compositional and processing spaces exhaustively. Incremental improvements in automation have accelerated particular laboratory tasks, yet they have not fully transformed the logic of discovery. A robot that performs a predefined protocol faster than a human still operates within a human-directed trial-and-error paradigm. The deeper opportunity lies in laboratories that can decide what to do next, execute experiments, learn from outcomes, and adapt their strategy without waiting for complete human redesign. This article proposes a conceptual framework for autonomous, closed-loop laboratories in materials discovery. The framework treats discovery as an adaptive, self-improving process rather than a linear sequence of isolated experiments. It argues that autonomous laboratories are not merely high-throughput platforms but emerging scientific systems in which artificial intelligence, robotics, characterization, data infrastructure, and human judgement are integrated through a shared feedback logic. The framework defines the closed-loop discovery cycle, identifies six conceptual layers of an autonomous laboratory, and explains how experiment selection, robotic execution, real-time data capture, and human oversight interact. It also highlights why data quality, machine-readability, and error handling are foundational rather than peripheral concerns. The resulting vision is of materials discovery as a learning ecosystem in which each experimental outcome improves not only the model of a material system but also the strategy by which future knowledge is generated.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2025 | Article: 41

A Multi-Fidelity Learning Framework for High-Entropy Alloy Design with Sparse DFT and Abundant CALPHAD Data
High-entropy alloy design requires exploring an enormous composition space in which conventional trial-and-error approaches are no longer viable. Density functional theory (DFT) delivers accurate formation energies and phase-stability predictions yet remains computationally prohibitive, yielding only sparse datasets of a few hundred structures per study. By contrast, the CALPHAD method furnishes abundant thermodynamic data across millions of compositions in seconds, yet it carries systematic biases when extrapolated beyond its binary and ternary calibration regimes. This conceptual framework presents a multi-fidelity learning strategy that systematically fuses sparse high-fidelity DFT data with abundant low-fidelity CALPHAD predictions to achieve near-DFT accuracy at CALPHAD-scale coverage. The framework rests on four tightly coupled components: a data-integration module that aligns CALPHAD and DFT outputs on identical compositions, a bias-correction module that learns the systematic mapping between the two fidelities, an uncertainty-propagation module that decomposes and combines fidelity-specific uncertainties, and an active-learning module that strategically selects the next DFT calculations where correction is most needed. An operational protocol translates these components into a repeatable workflow that begins with broad CALPHAD screening, proceeds through iterative DFT calibration, and converges when uncertainty falls below a designer-specified threshold. The approach reduces the DFT budget by approximately two orders of magnitude while preserving predictive fidelity, thereby opening previously inaccessible regions of high-entropy alloy space. Beyond immediate efficiency gains, the framework establishes a reusable blueprint for hybrid computational materials engineering in any system where abundant low-fidelity models coexist with sparse high-fidelity benchmarks. It therefore offers both a practical design pipeline for high-entropy alloys and a generalizable conceptual scaffold for multi-fidelity learning in complex concentrated alloys.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2025 | Article: 43

Perspective: From Prediction to Design — A Position on Optimization-Centric Workflows in Materials ML
Materials machine learning has achieved remarkable success in building accurate predictive models for properties such as formation energy, band gap, and mechanical strength. Yet the ultimate purpose of the field is not prediction but design: the discovery of new materials that deliver target properties under real-world constraints. This perspective argues that the community must shift from prediction-centric to optimization-centric workflows. Prediction-centric approaches focus on minimizing mean absolute error on randomly held-out test sets, while optimization-centric workflows aim to maximize (or minimize) a target property value within practical budgets, constraints, and sequential decision-making. These are fundamentally different objectives that demand different methods, evaluation metrics, and research priorities. A prediction-centric model may achieve low error on interpolation tasks yet fail catastrophically when asked to guide the search for extreme or out-of-distribution materials. In contrast, optimization-centric workflows treat the machine-learning model as a surrogate within an iterative loop that actively selects the next experiment or simulation. Four core principles underpin this shift: explicit goal specification that includes both the objective and all relevant constraints; robust constraint handling that distinguishes hard feasibility requirements from soft trade-offs; uncertainty awareness, where every prediction is accompanied by well-calibrated uncertainty estimates essential for balancing exploration and exploitation; and closed-loop integration that connects the model directly to automated or high-throughput experimentation and simulation. Adopting these principles will require new evaluation protocols that measure the best material discovered, sample efficiency, regret, constraint satisfaction, and extrapolation distance rather than isolated accuracy metrics. The implications extend to model development, benchmark design, journal standards, and funding priorities. Only by embracing optimization-centric workflows can materials machine learning fulfill its promise of accelerating discovery and delivering materials that solve pressing societal challenges.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2025 | Article: 48

Critical Assessment of Active Learning Benchmarks in Materials Discovery: Undisclosed Baselines and Optimistic Bias
Active learning has become a cornerstone strategy in data-driven materials discovery, promising to dramatically reduce the number of expensive simulations or experiments needed to identify high-performing materials. Proponents argue that uncertainty sampling, expected improvement, and other acquisition functions consistently outperform random selection, often by factors of 3–5× in iteration efficiency. Yet a closer examination of the benchmark studies published between 2017 and 2025 reveals a systematic pattern of optimistic bias that inflates these claims. This critical critique identifies five primary sources of overestimation: (1) weak or undisclosed random-sampling baselines, (2) unrealistic initial training sets that artificially favor active learning, (3) test-set leakage that prevents genuine extrapolation, (4) acquisition functions whose hyperparameters are implicitly tuned to the specific benchmark, and (5) incomplete reporting that hides variance and failure cases. Across the literature, random sampling is frequently presented as a naïve comparator yet proves surprisingly competitive once proper repetition, variance reporting, and realistic initial-set sizes are applied. Many studies fail to disclose the number of random seeds, the exact sampling distribution, or statistical significance tests, allowing small apparent gains to be reported as transformative. Initial training sets are often unrealistically small or already enriched with promising candidates, while test sets remain too similar to the training distribution, masking the true difficulty of exploration in vast chemical spaces. Acquisition functions are rarely subjected to hyperparameter robustness checks or evaluated on challenging out-of-distribution splits. The consequences extend beyond academic metrics: practitioners in industry and national laboratories risk deploying methods that underperform once transferred to real discovery campaigns. This critique, grounded exclusively in the 29 peer-reviewed studies listed in the reference section, calls for a new standard of rigor in active-learning evaluation. Only by adopting strong baselines, realistic initial conditions, extrapolation-aware test sets, full variance reporting, and public replication packages can the field move from optimistic benchmark theater to genuinely reliable acceleration of materials discovery.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2025 | Article: 54

Perspective: Autonomous Laboratories and Real-Time ML Feedback Loops — A Position for Closed-Loop Discovery
Materials discovery remains painfully slow. Traditional human-driven experimentation, followed by offline machine-learning analysis, requires weeks or months per iteration and leaves vast regions of chemical space unexplored. This position paper argues that autonomous laboratories equipped with real-time ML feedback loops represent not an incremental improvement but a necessary paradigm shift for the future of materials engineering. In these systems, robotic platforms handle synthesis and characterization while ML models continuously update and steer the next experiment, closing the discovery loop in hours rather than weeks. The current paradigm relies on human-in-the-loop decision-making, batch experimentation, and post-hoc ML training. Autonomous laboratories reverse this: robots execute synthesis and characterization tasks, a real-time ML engine analyzes streaming data, and an acquisition function immediately proposes the next candidate, all without human intervention for routine decisions. Early demonstrations have already shown accelerated discovery of battery electrolytes, perovskites, and catalysts. Real-time ML feedback loops demand online learning, rigorous uncertainty quantification, rapid acquisition functions, multi-objective optimization, constraint handling, human oversight for safety, and seamless data streaming. We articulate seven foundational principles for closed-loop discovery: integration-first design, speed as a first-class constraint, uncertainty-driven exploration, graceful degradation, data provenance, modularity, and open standards. These principles address the technical, operational, and cultural barriers that still prevent widespread adoption. While challenges remain—high initial costs, instrument integration, and long-duration experiments—the community now possesses the necessary ML maturity, robotic hardware, and orchestration tools to overcome them. This position calls for coordinated investment in shared autonomous-lab infrastructure, open standards, and training programs so that closed-loop discovery becomes the default workflow across academia and industry. Only then can materials science deliver the energy, sustainability, and electronics breakthroughs society urgently needs.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2026 | Article: 64

A Decade of Data-Driven Materials Engineering: Progress, Remaining Gaps, and Unlearned Lessons
Machine learning has transformed materials engineering through graph neural networks, equivariant architectures, generative models, and early autonomous laboratories. These advances enabled accurate property prediction, data-efficient force fields, and initial inverse design of inorganic crystals, supported by community databases such as the Materials Project and JARVIS. Generative frameworks now propose novel structures, while closed-loop platforms demonstrate early integration of prediction, synthesis, and robotic experimentation. However, major gaps remain: weak extrapolation beyond training distributions, inadequate handling of long-range interactions, poorly calibrated uncertainty quantification, limited synthesis prediction, underrepresentation of disordered materials, and insufficient experimental validation. Persistent unlearned lessons include benchmark biases that inflate generalization performance, poor reproducibility practices, misuse of invariant models for tensor properties, omission of random baselines in active learning, weak validity filtering in generative workflows, and near-absent failure reporting. This review critically assesses a decade of progress and shortcomings, grounded in 38 peer-reviewed publications. It highlights that hype around foundation models, generative inverse design, and autonomous labs often exceeds demonstrated impact. Realizing the field’s potential requires prioritizing extrapolation-focused architectures, mandatory experimental validation, shared failure registries, and stronger evidentiary standards to accelerate reliable discovery of functional materials.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 January 2026 | Article: 72

A Conceptual Framework for Uncertainty-Aware Active Learning in Compositional Design Spaces with Multiple Objectives
Compositional design spaces in materials engineering, particularly for alloys and multi-component systems, present exceptionally high-dimensional and sparsely populated landscapes that challenge conventional discovery workflows. A typical five-element alloy system sampled at 10% concentration increments can encompass millions of possible compositions, rendering exhaustive evaluation infeasible. Active learning has emerged as an essential paradigm for navigating these spaces efficiently, yet standard implementations often fail to address the simultaneous demands of multiple conflicting objectives—such as balancing mechanical strength against ductility, electrical conductivity against thermal stability, or performance against material cost—while properly accounting for uncertainty and practical constraints. This conceptual framework introduces an uncertainty-aware active learning approach tailored specifically for compositional design spaces with multiple objectives. The framework comprises four core components: robust uncertainty quantification that distinguishes epistemic from aleatoric contributions across heterogeneous composition regions, multi-objective acquisition functions grounded in Pareto optimization to explicitly explore trade-offs, systematic constraint handling for both hard thermodynamic limits and soft economic or manufacturability requirements, and adaptive sequential sampling strategies that prioritize informative candidates under limited evaluation budgets. The key innovation lies in a Pareto-based acquisition mechanism that integrates uncertainty estimates to balance exploration of uncertain regions, exploitation of promising trade-offs, and navigation of sparse high-dimensional manifolds. By embedding these elements into a unified conceptual structure, the framework provides design principles that enable more reliable, interpretable, and scalable materials discovery. It shifts the focus from single-objective optimization to holistic decision-making under uncertainty, ensuring that active learning not only accelerates identification of high-performing compositions but also respects real-world engineering constraints inherent to multi-component systems. The proposed approach is expected to guide both computational and experimental campaigns in alloy design, high-entropy materials, and mixed ionic conductors, ultimately fostering more sustainable and high-performance materials.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 January 2026 | Article: 73

Experimental Data Integration into ML Workflows for Real Alloys: A Review of Noisy, Sparse, and Heterogeneous Fusion
Integrating noisy, sparse, and heterogeneous experimental data with density functional theory (DFT) and machine learning (ML) is essential for reliable alloy design. While DFT enables high-throughput screening, its systematic biases limit predictive accuracy for real engineering alloys. This review synthesizes studies focused on fusing experimental measurements into ML–DFT workflows. We categorize five experimental data types—synthesis conditions, characterization data, property measurements, literature text, and industrial records—and identify six core challenges: noise, sparsity, heterogeneity, bias, missing metadata, and fragmentation. Six key integration strategies are examined: multi-fidelity learning, transfer learning, active learning with experimental feedback, Bayesian uncertainty modeling, multi-task learning, and physics-constrained augmentation. These approaches consistently reduce prediction errors by 30–60% compared to DFT-only models. Seven evidence-based best practices are distilled, emphasizing uncertainty reporting, data harmonization, experimental hold-out validation, and FAIR data sharing. Case studies demonstrate substantial gains in discovery efficiency for high-entropy alloys, superalloys, and phase diagrams. Remaining gaps, particularly the lack of standardized experimental databases and real-time feedback systems, are highlighted. This work provides a practical taxonomy and roadmap for developing experimentally grounded ML models that accelerate the design of high-performance alloys for aerospace, energy, and biomedical applications.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 January 2026 | Article: 75

Computational and Data-Driven Materials Engineering: High-Throughput Computational Screening Platforms, Workflows, and Discovery Outcomes
The field of computational and data-driven materials engineering has undergone rapid evolution, driven by advancements in high-throughput computational screening, machine learning algorithms, and integrated workflows that accelerate materials discovery. This review synthesizes recent developments in materials informatics, focusing on platforms that enable efficient exploration of vast chemical spaces through automated computations and data analytics. Key areas include the application of graph neural networks and representation learning for property prediction, active learning strategies to optimize experimental feedback loops, and the integration of multimodal datasets for enhanced model accuracy. High-throughput methods have facilitated discoveries in diverse domains, such as superconductors, battery materials, and high-entropy alloys, by combining density functional theory simulations with machine learning surrogates. Autonomous laboratories and closed-loop systems represent a paradigm shift, allowing self-driving experiments that minimize human intervention while maximizing discovery efficiency. Uncertainty quantification plays a critical role in guiding these processes, ensuring reliable predictions amid sparse data. This narrative review structures the landscape into computational ecosystems, workflow integrations, and discovery outcomes, highlighting cross-study synergies. It positions the field at the cusp of scalable, inverse design paradigms, where data-driven insights bridge simulation and experimentation to address grand challenges in materials science.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2023 | Article: 105

Computational and Data-Driven Materials Engineering: Multimodal Materials Datasets, Integration Frameworks, and Discovery Potential
The field of computational and data-driven materials engineering has transformed from traditional high-throughput simulations to sophisticated ecosystems integrating machine learning with multimodal datasets for accelerated discovery. This review synthesizes recent advancements in materials informatics, emphasizing the role of graph neural networks and deep learning in processing complex structural and property data. We examine multimodal datasets that combine experimental, computational, and textual modalities, enabling robust representation learning and uncertainty quantification. Integration frameworks are discussed, including active learning loops and multi-fidelity models that bridge simulation and experiment, addressing challenges like data sparsity and distribution shifts. The discovery potential is highlighted through applications in property prediction, inverse design, and autonomous systems, such as identifying stable alloys and energy materials. By providing an original synthesis of these elements, this article underscores the shift toward closed-loop workflows that enhance generalizability and interpretability, while identifying gaps in handling finite-temperature stability and disordered systems. Ultimately, these approaches promise to expand the known materials space by orders of magnitude, fostering innovations in sustainable technologies.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2023 | Article: 106

Transfer Learning in Computational Materials Engineering: Techniques and Case Studies
Transfer learning has become a cornerstone of computational materials engineering, addressing the fundamental tension between the exponential growth of high-throughput simulation data and the persistent scarcity of high-fidelity experimental labels. By repurposing knowledge encoded in large-scale computational repositories—ranging from density-functional theory (DFT) databases to molecular dynamics trajectories—transfer learning enables accurate property prediction, inverse design, and autonomous discovery even in data-constrained regimes. This review synthesizes the field’s maturation from early domain-adaptation approaches in microstructure informatics to contemporary foundation-model strategies that span inorganic crystals, organic polymers, and hybrid interfaces. We trace the evolution of techniques including graph-neural-network (GNN) pre-training, multi-fidelity fusion, and structure-aware fine-tuning, while highlighting their deployment in closed-loop pipelines that couple simulation with robotic experimentation. Case studies drawn from battery electrolytes, high-entropy alloys, and 2D heterostructures illustrate how hierarchical transfer frameworks achieve chemical accuracy with orders-of-magnitude fewer labels than scratch-trained models. The synthesis reveals a unifying computational workflow: pre-train on universal descriptors, adapt via frozen or low-rank updates, and close the loop through uncertainty-guided active learning. This infrastructure-level perspective underscores transfer learning’s role in transforming materials engineering from a trial-and-error discipline into a predictive, self-optimizing ecosystem.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 121

Governance Architectures for Self-Driving Laboratories in Computational Materials Engineering
The rapid evolution of computational and data-driven materials engineering has ushered in an era where self-driving laboratories (SDLs) promise to transform materials discovery by integrating automation, machine learning, and high-throughput experimentation into cohesive governance architectures. These architectures orchestrate the interplay between data generation, model training, and decision-making processes to enable closed-loop optimization in materials design. This review synthesizes recent advancements in SDL governance, focusing on how computational workflows—encompassing materials informatics, graph neural networks, representation learning, and uncertainty quantification—facilitate autonomous systems in addressing complex materials challenges. We examine the foundational elements of data-driven ecosystems, including multimodal datasets and simulation-experiment integration, and explore active learning strategies that balance exploration and exploitation in inverse design paradigms. Key governance components, such as orchestration platforms like ChemOS 2.0 and Bayesian active learning frameworks, are analyzed for their role in accelerating discovery cycles. By integrating perspectives from high-impact studies, we highlight how these architectures mitigate inefficiencies in traditional trial-and-error approaches, enabling scalable, reproducible materials innovation. The review positions SDL governance as a critical infrastructure for future materials engineering, emphasizing systems-level integration over isolated techniques. Ultimately, it underscores the potential of these architectures to democratize access to advanced materials development while identifying pathways for enhanced interoperability and robustness in computational ecosystems.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2025 | Article: 136
Filters
Clear All





Access type