Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Graph Neural Networks for Predicting Defect Formation Energies in 2D Materials
Two-dimensional (2D) materials have attracted considerable attention for next-generation electronic, optoelectronic, and catalytic applications; however, their performance is strongly influenced by the presence and stability of atomic-scale defects. Defect formation energy plays an essential role in defect prevalence, lattice stability, and functional behavior. Still, its evaluation remains challenging due to the complexity of defect-induced structural perturbations and the limitations of equilibrium-first-principles approaches. This paper presents an entirely conceptual framework that reframes defect formation energy estimation as a graph-structured inference problem. Leveraging graph neural networks (GNNs), the proposed defect-aware graph neural architecture (DAGNA) represents pristine and defect-perturbed lattices as coupled relational graphs, enabling structured propagation of defect-induced information across spatial scales. Instead of proposing a predictive or validated model, the framework explains how hierarchical message passing, defect-aware embeddings, and physics-constrained aggregation can be organized to regulate information flow under defect perturbations in two-dimensional systems. By synthesizing advances in graph theory, the physics of defects, and materials-focused AI, this work provides an operational decision-making framework for reasoning about defect formation energy without relying on empirical datasets or simulations. This framework contributes to the theoretical foundations of applied artificial intelligence in materials science. It provides a clear, physically grounded architecture for future studies in defect-aware materials modeling and defect engineering.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2023 | Article: 35

Generalization in Materials AI: A Theoretical Distinction between New Compositions, New Structures, and New Physics
The integration of artificial intelligence into materials science has accelerated property prediction and high-throughput screening. Yet, the field’s progress hinges on models’ ability to generalize beyond their training distributions. Existing literature often addresses generalization in broad terms, focusing on out-of-distribution performance or extrapolation without distinguishing the qualitative nature of material novelty. This conceptual manuscript introduces a novel theoretical framework for categorizing generalization in materials AI into three distinct levels: new compositions (variations within known structural families), new structures (alternative atomic arrangements or topologies), and new physics (emergence of phenomena governed by mechanisms absent from the training data). Drawing on recent advances in graph neural networks, scalable deep learning, and materials representations, we synthesize evidence that current models achieve reasonable interpolation within familiar domains but encounter progressively greater difficulties across these levels. The proposed distinction provides a structured lens for analyzing model limitations, interpreting benchmark results, and guiding the design of future architectures and training strategies. By formalizing these categories, the framework aims to advance theoretical understanding of generalization in materials AI, emphasizing the need for targeted approaches at each level to enable reliable discovery of novel materials.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 55

The Attention Economy of Materials AI: How Model Focus Shapes Scientific Attention Allocation
In the expanding domain of artificial intelligence applied to materials science, computational models do not merely predict properties or accelerate screening; they function as subtle but powerful mechanisms that allocate finite scientific attention across an effectively infinite chemical space. By prioritizing certain compositional regions, structural motifs, or property axes while de-emphasizing others, these systems implicitly decide which questions will be asked, which hypotheses will be tested, and which materials classes will receive downstream experimental or theoretical investment. This position paper argues that Materials AI operates as an attention-allocation infrastructure whose architectural choices reshape the trajectory of discovery itself, transforming what was once an open-ended scientific exploration into a directed economy of focus. Drawing on the well-established “attention economy” metaphor from information systems and cognitive science, we introduce the parallel concept of scientific attention capital—the limited pool of researcher time, funding, instrumentation access, and collective curiosity that models now mediate and, in many cases, ration. Rather than viewing model-induced focus as a neutral technical artifact, we distinguish productive attention (focused investment that yields rapid, high-impact advances in targeted domains) from pathological attention (self-reinforcing loops that create blind spots, reward hacking, and representational injustice). The perspective developed here suggests that recognizing Materials AI as an attention-shaping force carries immediate implications for how the community designs and audits. It deploys these systems if the goal is to preserve the generative openness that has historically driven materials innovation. Ultimately, treating attention allocation as an explicit design variable rather than an incidental byproduct offers a conceptual framework for ensuring that the next generation of Materials AI expands, rather than contracts, the horizons of scientific possibility.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 96

The Problem of Convergent Scientific Narratives in Single-Model Materials Regimes
Materials AI is rapidly converging toward single-model regimes in which a handful of dominant architectures, particularly graph neural networks, have become the de facto standard for property prediction, inverse design, and materials discovery. This model monoculture does not merely reflect technical superiority; it actively produces convergent scientific narratives that shape what the community considers valid knowledge, worthwhile problems, and genuine progress in the field. The present critique identifies four interlocking epistemic risks of this convergence: epistemic narrowing, suppression of alternatives, paradigm lock-in, and the illusion of consensus. These risks threaten the long-term robustness of materials science by limiting the diversity of phenomena that can be observed, the range of methods that can be explored, and the kinds of disagreement that can be productively acknowledged. The consequences include missed discoveries in complex materials systems, methodological stagnation, overconfidence in model outputs, and path-dependent research trajectories that will prove difficult to reverse. Alternative approaches grounded in deliberative methodological pluralism, adversarial benchmarking, narrative diversity, paradigm auditing, and deliberate switching-cost reduction are therefore proposed as necessary correctives if the field is to preserve its epistemic openness while retaining the undeniable benefits of data-driven methods.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2026 | Article: 150

From High-Throughput Computation to Autonomous Discovery: A Review of Closed-Loop Data Infrastructures in Materials Engineering
The field of materials engineering has undergone a profound transformation through the integration of high-throughput computation and data-driven methodologies, evolving from traditional trial-and-error approaches to sophisticated closed-loop systems that accelerate discovery. This review synthesizes recent advancements in computational and data-driven materials ecosystems, focusing on the infrastructure enabling autonomous discovery. Key elements include materials informatics platforms that leverage machine learning for property prediction and inverse design, graph neural networks for representation learning, and high-throughput computational workflows that generate multimodal datasets. We examine the progression from static high-throughput screening to dynamic, closed-loop paradigms incorporating active learning, uncertainty quantification, and simulation-experiment integration. Autonomous laboratories represent a pinnacle of this evolution, where AI orchestrates iterative cycles of hypothesis generation, experimentation, and refinement. The synthesis highlights how these infrastructures bridge computational predictions with experimental validation, fostering inverse materials design and optimizing resource allocation in complex chemical spaces. Challenges in data interoperability and model generalizability are noted, alongside prospects for scalable, self-optimizing systems. Overall, this review positions closed-loop data infrastructures as foundational to next-generation materials engineering, promising accelerated innovation in areas like energy storage, catalysis, and structural materials. By integrating diverse literature, we provide a systems-level perspective on how these tools are reshaping the discovery landscape.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 March 2022 | Article: 81

Representation Learning in Materials Science: Architectures, Data Modalities, and Discovery Applications
The field of materials science has undergone a transformative shift with the integration of computational and data-driven approaches, particularly through representation learning techniques that enable efficient handling of complex materials data. This review synthesizes recent advancements in architectures for representation learning, encompassing graph neural networks, attention-based models, and physics-inspired embeddings, which facilitate the extraction of meaningful features from diverse data modalities such as atomic structures, stoichiometries, and spectroscopic data. By bridging traditional computational methods with machine learning, these representations have accelerated property prediction, inverse design, and materials discovery applications, addressing challenges in high-dimensional spaces and sparse datasets. The scope of this narrative review covers the evolution from basic informatics to sophisticated multimodal integrations, highlighting how data ecosystems and learning frameworks contribute to autonomous discovery pipelines. A systems-level perspective is adopted to integrate cross-study insights, revealing synergies between representation learning and closed-loop systems that couple simulations with experiments. Looking ahead, the review posits that continued refinement of these architectures will drive scalable, AI-guided materials engineering, fostering innovations in energy, electronics, and structural materials while emphasizing the need for robust, interpretable models in real-world applications.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 March 2022 | Article: 82

Graph Neural Networks for Materials Property Prediction: A Decadal Review of Advances and Limits
The advent of graph neural networks (GNNs) has revolutionized computational materials engineering by enabling sophisticated representations of atomic structures and interactions for property prediction. This review synthesizes key developments in GNN architectures tailored for materials science, focusing on their application in predicting mechanical, electronic, and thermodynamic properties of diverse materials systems, including polycrystals, metal-organic frameworks, and perovskites. Drawing from high-impact studies, we examine the evolution from basic crystal graph convolutional networks to advanced variants incorporating transfer learning, data augmentation, and force field integration. The synthesis highlights how GNNs address challenges in materials data sparsity and structural complexity through graph-based featurization, leading to improved accuracy in property forecasts compared to traditional machine learning methods. We integrate perspectives on GNNs' role in broader data-driven ecosystems, including their synergy with active learning for autonomous discovery pipelines. Limitations such as interpretability and scalability are critically assessed, alongside advances in benchmark frameworks that standardize evaluations. The review positions GNNs as a cornerstone of next-generation materials informatics, accelerating the design of high-performance materials for energy, catalysis, and structural applications. Future outlooks emphasize hybrid integrations with physics-based simulations to bridge experimental and computational gaps, fostering closed-loop systems for rapid materials innovation. This narrative underscores the transformative potential of GNNs in reshaping materials engineering paradigms.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2022 | Article: 84

Latent Space Collapse in Materials Representation Learning: A Systems-Level Conceptual Analysis
In the evolving landscape of computational materials engineering, data-driven approaches have transformed traditional discovery paradigms by integrating machine learning with high-throughput simulations and experimental workflows. Representation learning, particularly through graph neural networks and deep architectures, enables the encoding of complex material structures into latent spaces that facilitate property prediction, inverse design, and autonomous exploration. However, latent space collapse—where embeddings fail to preserve structural diversity or physicochemical distinctions—poses a systemic challenge, undermining the reliability of inference in materials informatics ecosystems. This conceptual analysis frames latent space collapse as an emergent property of interconnected data infrastructures, model architectures, and discovery pipelines, drawing on systems-level interactions across multimodal datasets and uncertainty quantification mechanisms. We introduce the Representation Integrity Framework (RIF), a novel interpretive structure that dissects collapse dynamics through layers of data encoding, model compression, and feedback-driven steering. By examining computational trade-offs and epistemic risks, RIF highlights pathways for resilient representation learning, such as enhanced multimodal integration and adaptive uncertainty handling. Implications extend to closed-loop systems, where mitigating collapse could optimize simulation-experiment coupling and accelerate inverse materials design. This framework advances a balanced view of AI in materials science, emphasizing infrastructure resilience over isolated algorithmic fixes, and informs future developments in foundation models for scientific discovery.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2022 | Article: 88

Topology without Physics: Structural Abstraction Limits in Graph-Based Materials Models
The advent of computational and data-driven approaches in materials engineering has transformed discovery pipelines, leveraging machine learning and graph-based representations to navigate vast chemical spaces. However, these models often prioritize topological abstractions over intrinsic physical mechanisms, leading to epistemic constraints in predictive accuracy and interpretability. This manuscript introduces a conceptual framework that dissects the structural abstraction limits inherent in graph-based materials models, emphasizing the trade-offs between computational efficiency and physical fidelity. By synthesizing insights from materials informatics and representation learning, we explore how graph neural networks decouple topological features from underlying physics, potentially hindering autonomous discovery systems and inverse design workflows. The framework delineates layers of abstraction, from data ingestion to inference, highlighting feedback loops that amplify abstraction-induced uncertainties. Implications extend to high-throughput computation, multimodal datasets, and uncertainty quantification, advocating for integrated infrastructures that balance abstraction with mechanistic reintegration. This analysis fosters a deeper understanding of computational steering in materials AI, guiding future developments toward more robust, physics-aware discovery paradigms without empirical validation. Ultimately, addressing these limits could enhance the reliability of data-driven materials engineering ecosystems.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2022 | Article: 92

Uncertainty as Infrastructure: Governing Confidence in Data-Driven Materials Pipelines
The field of computational and data-driven materials engineering has transformed traditional discovery processes through the integration of machine learning, high-throughput computations, and autonomous systems. However, as these pipelines scale, the management of uncertainty emerges as a foundational infrastructure rather than a mere analytical byproduct. This manuscript conceptualizes uncertainty not as an obstacle but as an enabling framework for governing confidence in materials informatics workflows. By synthesizing recent advancements in representation learning, graph neural networks, and uncertainty quantification, we identify epistemic gaps in current data-driven ecosystems, where confidence in predictions often remains opaque or inadequately integrated into discovery loops. We introduce the Confidence Governance Framework (CGF), a layered conceptual architecture that embeds uncertainty quantification as a core infrastructural element, facilitating dynamic interactions between data representations, model inferences, and discovery steering. This framework emphasizes computational trade-offs in multimodal datasets and simulation-experiment couplings, promoting robust, interpretable pipelines. Implications extend to enhanced autonomy in inverse design and closed-loop experimentation, fostering resilient materials engineering paradigms. Through this lens, uncertainty becomes a strategic asset for calibrating epistemic risks and optimizing resource allocation in AI-assisted materials research.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2022 | Article: 93

Multimodal, Physics-Informed Machine Learning for Accelerated Materials Design and Discovery
In the evolving landscape of computational materials engineering, the integration of multimodal data sources with physics-informed machine learning paradigms promises to revolutionize the pace and precision of materials design and discovery. This conceptual manuscript explores the synergies between diverse data modalities—ranging from experimental spectra to simulation-derived properties—and machine learning models constrained by physical laws, aiming to address persistent challenges in data scarcity, model generalizability, and discovery efficiency within materials science. By synthesizing recent advancements in representation learning, graph neural networks, and autonomous systems, we identify a conceptual gap in holistic frameworks that unify multimodal inputs with physics-based priors for accelerated inverse design. We introduce a novel conceptual framework, termed the Multimodal Physics-Constrained Discovery Engine (MPCDE), which structures data-model-discovery pipelines through layered interactions, feedback mechanisms, and epistemic steering logics. This framework emphasizes computational workflows that balance representation fidelity with inference robustness, incorporating uncertainty quantification to mitigate risks in high-throughput settings. Implications for the field include enhanced coupling of simulation and experimentation, improved scalability of foundation models, and streamlined closed-loop discovery systems. Ultimately, this work posits interpretive insights into how such integrated approaches can transform materials informatics into a more predictive and autonomous discipline, fostering innovations in energy, electronics, and structural materials.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2023 | Article: 100

Simulation Priors in Machine Learning Materials Models: Hidden Physics Assumptions
The integration of machine learning into materials engineering has transformed discovery pipelines by leveraging vast simulation-generated datasets and high-throughput computational workflows. Within this data-driven paradigm, models frequently incorporate simulation priors—implicit assumptions derived from physical approximations, boundary conditions, and discretization choices embedded in first-principles calculations or molecular dynamics trajectories. These priors, often hidden within representation learning and graph-based architectures, introduce epistemic biases that propagate through inference to downstream tasks such as inverse design and closed-loop experimentation. A key conceptual gap lies in the lack of systematic frameworks for articulating and managing these assumptions as integral components of the computational infrastructure rather than incidental data artifacts. This article introduces the Simulation Prior Articulation Framework (SPAF), an original systems-level conceptual structure that delineates layered processing of multimodal materials data, explicit prior extraction from simulation ecosystems, integration into deep learning architectures, and steering of discovery pipelines via feedback mechanisms. SPAF emphasizes representation–inference interactions, computational workflow dynamics, and infrastructure trade-offs to enhance simulation–experiment coupling without empirical benchmarking. By framing hidden physics assumptions as addressable epistemic structures, the framework provides integrative insights for materials informatics, foundation models, and autonomous discovery systems, supporting more transparent and robust data-driven materials engineering pipelines.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2023 | Article: 103

The Interpretability–Complexity Paradox in Deep Materials Networks
In the evolving landscape of computational and data-driven materials engineering, deep neural networks have emerged as powerful tools for accelerating materials discovery and design. These architectures leverage vast multimodal datasets, high-throughput computations, and representation learning to model complex structure-property relationships in materials systems. However, a fundamental tension arises: as network complexity increases to capture intricate physical phenomena, interpretability diminishes, hindering the extraction of scientific insights essential for advancing materials informatics. This interpretability-complexity paradox poses a significant barrier to integrating deep models into autonomous discovery pipelines, where uncertainty quantification and simulation-experiment coupling demand transparent decision-making. To address this gap, we introduce the Interpretive Complexity Equilibrium Framework (ICEFrame), a novel conceptual structure that conceptualizes the dynamic interplay between model depth, representational fidelity, and epistemic transparency in deep materials networks. ICEFrame delineates layered interactions across data ingestion, architectural scaling, and inference steering, incorporating feedback mechanisms to balance trade-offs without empirical validation. This framework offers interpretive lenses for navigating complexity in graph neural networks and foundation models for science, fostering more robust closed-loop experimentation and inverse design strategies. By reframing the paradox through systems-level insights, ICEFrame implications extend to enhancing discovery steering logics in materials AI, ultimately promoting sustainable innovation in computational materials ecosystems.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2023 | Article: 104

Computational and Data-Driven Materials Engineering: Multimodal Materials Datasets, Integration Frameworks, and Discovery Potential
The field of computational and data-driven materials engineering has transformed from traditional high-throughput simulations to sophisticated ecosystems integrating machine learning with multimodal datasets for accelerated discovery. This review synthesizes recent advancements in materials informatics, emphasizing the role of graph neural networks and deep learning in processing complex structural and property data. We examine multimodal datasets that combine experimental, computational, and textual modalities, enabling robust representation learning and uncertainty quantification. Integration frameworks are discussed, including active learning loops and multi-fidelity models that bridge simulation and experiment, addressing challenges like data sparsity and distribution shifts. The discovery potential is highlighted through applications in property prediction, inverse design, and autonomous systems, such as identifying stable alloys and energy materials. By providing an original synthesis of these elements, this article underscores the shift toward closed-loop workflows that enhance generalizability and interpretability, while identifying gaps in handling finite-temperature stability and disordered systems. Ultimately, these approaches promise to expand the known materials space by orders of magnitude, fostering innovations in sustainable technologies.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2023 | Article: 106

Data-Driven Materials Engineering: Inverse Design Strategies, Machine Learning Architectures, and Application Domains
The advent of data-driven approaches has revolutionized materials engineering, enabling inverse design strategies that prioritize target properties to guide material synthesis and optimization. This review synthesizes recent advancements in machine learning architectures tailored for materials informatics, including graph neural networks and representation learning frameworks that capture atomic-scale interactions and multiscale phenomena. We examine the integration of high-throughput computations with experimental workflows, highlighting closed-loop systems that incorporate active learning and uncertainty quantification to accelerate discovery. Key application domains span energy materials, metamaterials, and catalytic systems, where multimodal datasets facilitate simulation-experiment synergies. By analyzing computational ecosystems, we underscore the shift from forward modeling to inverse paradigms, emphasizing autonomous laboratories that iteratively refine hypotheses through data feedback loops. Challenges in generalizability and data scarcity are contextualized within broader systems integration, offering a cohesive perspective on how these tools reshape materials design. This narrative integrates cross-study insights to propose unified frameworks for scalable, data-centric engineering, bridging theoretical models with practical implementations in computational materials science.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2023 | Article: 107

Closed-World Training in an Open Materials Universe
In the rapidly evolving field of computational and data-driven materials engineering, machine learning models are increasingly trained on curated datasets that represent a closed-world approximation of material properties and behaviors. However, the broader materials universe encompasses vast, unexplored compositional spaces, dynamic environmental interactions, and emergent phenomena that defy static boundaries. This conceptual manuscript addresses the inherent tension between closed-world training paradigms—characterized by finite, labeled data regimes—and the open, infinite nature of materials discovery. We introduce a novel conceptual framework, termed the Adaptive Boundary Inference Architecture (ABIA), which integrates representation learning, uncertainty-aware feedback mechanisms, and multi-scale inference logics to navigate this disparity. ABIA conceptualizes training as a dynamic process where model boundaries adapt through iterative interactions between data representations and discovery pipelines, fostering resilience to out-of-distribution materials. By synthesizing recent advances in graph neural networks, foundation models, and autonomous systems, the framework highlights computational steering strategies that balance exploitation of known data with exploration of open spaces. Implications extend to enhanced inverse design, multimodal integration, and epistemic risk management in materials informatics, ultimately advancing sustainable and efficient materials engineering workflows. This work underscores the need for interpretive systems that transcend traditional closed-loop constraints, promoting a more holistic approach to data-driven discovery in an unbounded materials landscape.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2024 | Article: 109

Knowledge Graphs vs Property Predictors: Competing Infrastructures of Materials Intelligence
In the evolving landscape of computational and data-driven materials engineering, the integration of machine learning techniques has transformed traditional discovery paradigms into intelligent, autonomous systems. Materials informatics leverages vast datasets from high-throughput computations and multimodal sources to accelerate the design of novel materials with tailored properties. However, a conceptual gap persists in understanding the infrastructural roles of knowledge graphs and property predictors as competing yet complementary architectures for materials intelligence. Knowledge graphs offer relational representations that capture complex interdependencies among materials entities, enabling semantic querying and inference across disparate data modalities. In contrast, property predictors, often based on graph neural networks or deep learning models, focus on direct regression or classification of material attributes, prioritizing predictive accuracy over holistic system integration. This manuscript introduces a novel conceptual framework, termed the Dual-Infrastructure Materials Cognition (DIMC) model, which interprets the dynamic interplay between these infrastructures through layered computational workflows and feedback mechanisms. By examining representation learning, uncertainty quantification, and closed-loop discovery logics, the framework elucidates trade-offs in scalability, interpretability, and epistemic robustness. Implications for the field include enhanced steering of autonomous discovery systems, improved coupling of simulation and experimentation, and refined strategies for inverse materials design. Ultimately, this interpretive lens fosters a more cohesive ecosystem for materials intelligence, bridging isolated predictive tools with knowledge-centric infrastructures to advance data-driven innovation in materials science.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2024 | Article: 113

Latent Anisotropy: Directional Bias in Materials Embedding Spaces
In the evolving landscape of computational and data-driven materials engineering, embedding spaces serve as foundational representations that encode material properties, structures, and behaviors into vectorial forms amenable to machine learning workflows. These spaces facilitate high-throughput screening, inverse design, and autonomous discovery by bridging atomic-scale simulations with macroscopic predictions. However, inherent directional biases—termed latent anisotropy—emerge from the interplay of data modalities, architectural choices in neural networks, and inference dynamics, potentially skewing discovery pathways toward certain material classes or property regimes. This conceptual manuscript identifies a critical gap in understanding how such biases propagate through materials informatics pipelines, influencing the epistemic reliability of AI-assisted materials research. We introduce the Anisotropic Representation Cascade (ARC) framework, which conceptualizes embedding spaces as multi-layered systems where directional preferences arise from representation encoding, propagation through graph-based architectures, and feedback in closed-loop systems. By integrating insights from uncertainty quantification and multimodal data fusion, ARC elucidates trade-offs in computational steering logics that balance exploration breadth with directional fidelity. Implications extend to enhancing robustness in foundation models for materials science, fostering more equitable navigation of chemical spaces, and informing infrastructure designs that mitigate bias amplification in simulation-experiment couplings. This work underscores the need for interpretive tools in data-driven paradigms to ensure unbiased acceleration of materials innovation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2024 | Article: 114

Pretraining on Matter: Conceptual Limits of Foundation Models for Materials Science
The advent of foundation models, large-scale pre-trained architectures adapted from natural language processing paradigms, has permeated computational materials science, promising accelerated discovery through data-driven inference. In materials engineering, these models leverage multimodal datasets encompassing atomic structures, properties, and simulations to enable representation learning across scales. However, inherent conceptual limits arise from the interplay between materials' physical hierarchies—spanning quantum to macroscopic levels—and the inductive biases embedded in pretraining strategies. This manuscript synthesizes recent advancements in machine learning architectures, such as graph neural networks and multimodal integration, within materials informatics ecosystems. It identifies epistemic boundaries where foundation models falter in capturing causality, uncertainty, and domain-specific invariances, potentially leading to misaligned discovery pipelines. To address these, we introduce the Matter Pretraining Boundary Framework (MPBF), a conceptual architecture that delineates layers of data assimilation, representational abstraction, and inference steering to mitigate limits in autonomous materials design. Implications extend to high-throughput computation, inverse design, and simulation-experiment coupling, fostering more robust computational workflows in materials engineering. By interpreting these limits through systems-level dynamics, the framework guides infrastructure trade-offs, enhancing the reliability of data-driven paradigms without empirical validation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 115

Representation Drift in Iterative Materials Learning Systems
In the evolving landscape of computational and data-driven materials engineering, iterative learning systems have become pivotal for accelerating materials discovery through integrated machine learning pipelines and high-throughput computations. These systems, encompassing active learning loops and closed-loop experimentation, rely on dynamic representations of materials properties and structures to guide successive iterations of model refinement and data acquisition. However, a critical yet underexplored phenomenon emerges: representation drift, where iterative updates inadvertently alter the semantic fidelity of learned embeddings, potentially leading to misaligned inferences across discovery cycles. This conceptual manuscript identifies this gap within materials informatics ecosystems, highlighting how drift manifests in graph neural networks, multimodal datasets, and uncertainty-aware frameworks. To address this, we introduce the Iterative Representation Stabilization Framework (IRSF), a novel conceptual architecture that integrates stabilization mechanisms across data ingestion, model adaptation, and inference steering layers. IRSF conceptualizes drift as a systemic interaction between feedback loops and representation spaces, offering interpretive insights into maintaining epistemic consistency in autonomous discovery workflows. Implications extend to enhancing the robustness of foundation models for science, simulation-experiment couplings, and inverse design paradigms, fostering more reliable computational steering in materials engineering. By framing representation drift through infrastructure-level trade-offs, this work provides a foundational lens for interpreting iterative dynamics, ultimately supporting sustainable advancements in data-driven materials paradigms.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 116

Foundation Models in Materials Science: Emerging Architectures and Training Paradigms
The convergence of large-scale machine learning with materials engineering is reshaping how new materials are conceived, predicted, and realized. Foundation models—pre-trained architectures that learn generalizable, multimodal representations from expansive datasets—are emerging as the computational backbone for next-generation discovery pipelines. This narrative review synthesizes the computational and data-driven ecosystems that have enabled their rise, drawing on advances in materials informatics, graph-based representation learning, and autonomous experimentation. We trace the progression from early machine learning applications in property prediction to scalable graph neural networks that capture atomic-scale interactions with unprecedented fidelity. High-throughput computation and multimodal data integration have created the knowledge bases necessary for training models that generalize across chemical spaces. Central to this evolution are closed-loop systems, where foundation-like models orchestrate active learning, uncertainty-aware selection, and seamless simulation–experiment feedback. Through an original integrative analysis, we identify recurring architectural principles—such as hierarchical graph convolutions, contrastive pre-training, and multi-task optimization—and training paradigms that balance exploration with exploitation in vast design spaces. These elements collectively address longstanding bottlenecks in inverse design, property optimization, and length-scale bridging. Positioned at the interface of computational infrastructure and autonomous discovery, this review provides a systems-level perspective on how foundation models are poised to compress the materials innovation timeline from decades to months, while maintaining rigorous physical grounding.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 119

Knowledge Graphs for Materials Discovery: Data Structuring, Reasoning, and Applications
Knowledge graphs (KGs) have emerged as a pivotal infrastructure in computational and data-driven materials engineering, enabling structured representation, reasoning, and integration of heterogeneous data for accelerated discovery. By organizing materials data into interconnected entities and relationships, KGs facilitate advanced querying, inference, and machine learning applications across domains such as materials informatics, high-throughput computation, and inverse design. This review synthesizes recent advancements in KG construction from multimodal datasets, including text corpora, biomolecular integrations, and crystalline structures. We examine how graph neural networks and representation learning enhance molecular contrastive learning and pre-training frameworks for improved molecular representations. In the landscape of computational materials ecosystems, KGs support semantic integration and terminology standardization, bridging simulation and experiment through active learning systems and uncertainty quantification. Applications in autonomous laboratories highlight closed-loop discovery, where KGs enable dynamic knowledge propagation and event-sourced provenance management. We provide an original synthesis framing KGs as unifying backbones for data-model-experiment cycles, emphasizing systems-level integration over isolated tools. Challenges in scalability and interoperability are noted, with future directions toward hybrid human-AI workflows. This narrative underscores KGs' role in transforming materials discovery from empirical to predictive paradigms, fostering interdisciplinary convergence in materials science.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 120

Transfer Learning in Computational Materials Engineering: Techniques and Case Studies
Transfer learning has become a cornerstone of computational materials engineering, addressing the fundamental tension between the exponential growth of high-throughput simulation data and the persistent scarcity of high-fidelity experimental labels. By repurposing knowledge encoded in large-scale computational repositories—ranging from density-functional theory (DFT) databases to molecular dynamics trajectories—transfer learning enables accurate property prediction, inverse design, and autonomous discovery even in data-constrained regimes. This review synthesizes the field’s maturation from early domain-adaptation approaches in microstructure informatics to contemporary foundation-model strategies that span inorganic crystals, organic polymers, and hybrid interfaces. We trace the evolution of techniques including graph-neural-network (GNN) pre-training, multi-fidelity fusion, and structure-aware fine-tuning, while highlighting their deployment in closed-loop pipelines that couple simulation with robotic experimentation. Case studies drawn from battery electrolytes, high-entropy alloys, and 2D heterostructures illustrate how hierarchical transfer frameworks achieve chemical accuracy with orders-of-magnitude fewer labels than scratch-trained models. The synthesis reveals a unifying computational workflow: pre-train on universal descriptors, adapt via frozen or low-rank updates, and close the loop through uncertainty-guided active learning. This infrastructure-level perspective underscores transfer learning’s role in transforming materials engineering from a trial-and-error discipline into a predictive, self-optimizing ecosystem.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 September 2024 | Article: 121

Governance Architectures for Self-Driving Laboratories in Computational Materials Engineering
The rapid evolution of computational and data-driven materials engineering has ushered in an era where self-driving laboratories (SDLs) promise to transform materials discovery by integrating automation, machine learning, and high-throughput experimentation into cohesive governance architectures. These architectures orchestrate the interplay between data generation, model training, and decision-making processes to enable closed-loop optimization in materials design. This review synthesizes recent advancements in SDL governance, focusing on how computational workflows—encompassing materials informatics, graph neural networks, representation learning, and uncertainty quantification—facilitate autonomous systems in addressing complex materials challenges. We examine the foundational elements of data-driven ecosystems, including multimodal datasets and simulation-experiment integration, and explore active learning strategies that balance exploration and exploitation in inverse design paradigms. Key governance components, such as orchestration platforms like ChemOS 2.0 and Bayesian active learning frameworks, are analyzed for their role in accelerating discovery cycles. By integrating perspectives from high-impact studies, we highlight how these architectures mitigate inefficiencies in traditional trial-and-error approaches, enabling scalable, reproducible materials innovation. The review positions SDL governance as a critical infrastructure for future materials engineering, emphasizing systems-level integration over isolated techniques. Ultimately, it underscores the potential of these architectures to democratize access to advanced materials development while identifying pathways for enhanced interoperability and robustness in computational ecosystems.
Journal of Computational and Data-Driven Materials Engineering
Review | Open access | 18 September 2025 | Article: 136
Filters
Clear All





Access type