Materials informatics has emerged as a central paradigm in contemporary materials science, leveraging machine learning and data-driven modeling to accelerate materials discovery, optimization, and deployment. Despite substantial advances in predictive accuracy, most existing approaches remain fundamentally correlational, limiting their reliability under distribution shifts, experimental interventions, and real-world deployment scenarios. This reliance on correlation constrains scientific interpretability and undermines the capacity of AI systems to function as genuine instruments of materials reasoning. Causality offers a principled framework for overcoming these limitations by explicitly modeling cause-and-effect relationships among composition, processing, structure, and properties. This narrative review synthesizes conceptual progress in integrating causal inference into materials informatics, examining foundational causal frameworks, advances in causal discovery, and hybrid causal–machine learning approaches, and emerging applications across materials domains such as nanocatalysis, ferroelectrics, and electrochemical energy storage. We critically analyze persistent challenges—including data scarcity, assumption violations, limited external validity, and computational and epistemic constraints—that currently hinder widespread adoption. Drawing exclusively on peer-reviewed literature published, the review emphasizes thematic and epistemic developments rather than algorithmic prescriptions. We argue that causality represents a structural shift in how AI systems contribute to materials science: from correlational predictors to intervention-aware, mechanism-aligned reasoning tools. By articulating future directions centered on hybrid modeling, domain-knowledge integration, and interdisciplinary collaboration, this review positions causality as a necessary foundation for robust, generalizable, and scientifically legitimate materials informatics.
This review systematically examines the conceptual treatment of causality within materials informatics literature published between 2017 and 2024, drawing exclusively on a curated set of 26 studies identified through targeted and broadened searches across Web of Science, Scopus, arXiv, and specialized databases using terms such as “causal inference,” “causality materials informatics,” “structural causal model,” “directed acyclic graph,” “intervention materials design,” and “counterfactual materials prediction,” with inclusion criteria focused on relevance to materials AI while allowing broader engineering and general causal frameworks where they intersect with materials problems. The analysis reveals a pronounced dominance of correlation-based approaches in materials artificial intelligence, where predictive models achieve impressive statistical fits for structure-property relationships yet seldom progress to robust causal claims, as evidenced by the majority of surveyed works prioritizing accuracy metrics over interventional or counterfactual reasoning. Key causal concepts and frameworks, primarily drawn from Pearl’s foundational hierarchy of association, intervention, and counterfactuals as well as structural causal models and directed acyclic graphs, are introduced and contrasted with their limited adoption in the field. Causal methods that have been applied, albeit sparingly, to materials informatics—ranging from data-driven causal discovery to Bayesian causal modeling—are surveyed alongside their strengths and context-specific limitations. Persistent challenges, including the rarity of randomized interventions in experimental materials workflows and the confounding effects inherent in high-dimensional observational datasets, are highlighted as barriers that leave substantial gaps in the literature. Ultimately, this review offers targeted recommendations for authors, reviewers, and the broader community to integrate causal reasoning more explicitly, thereby moving materials informatics from correlational prediction toward actionable intervention and counterfactual understanding essential for autonomous materials design.
In the evolving landscape of computational and data-driven materials engineering, machine learning techniques have revolutionized the discovery and optimization of materials by leveraging vast datasets to identify patterns and correlations. However, this reliance on correlation-driven approaches often overlooks the underlying causal mechanisms that govern material properties and behaviors, leading to inherent limitations in the generalizability and robustness of designed materials. This manuscript explores the conceptual boundaries of optimization strategies that prioritize statistical associations over causal understanding within materials informatics ecosystems. We introduce a novel conceptual framework, termed the Correlation Boundary Architecture (CBA), which delineates the epistemic constraints imposed by correlation-centric pipelines in materials design. The CBA integrates representation learning, inference dynamics, and feedback structures to highlight how data-driven optimizations can falter in extrapolative scenarios, such as novel chemical spaces or extreme conditions. By synthesizing recent advancements in graph neural networks, high-throughput computations, and uncertainty quantification, we articulate the trade-offs between computational efficiency and causal fidelity. Implications extend to autonomous discovery systems and inverse design paradigms, suggesting pathways for hybrid frameworks that mitigate correlation biases through enhanced interpretive layers. This work underscores the need for computational steering logics that balance correlative power with causal awareness, fostering more resilient materials engineering practices.