This review article examines the literature on scientific explanation in AI-driven materials science, focusing on the conceptual foundations and evaluative criteria that distinguish genuine scientific explanation from the predictive and interpretive outputs commonly produced by machine learning models in the field. The methodology involved a systematic search across major databases and targeted journals using predefined strings related to scientific explanation, explainable AI (XAI), and interpretability in materials contexts, resulting in the inclusion of 30 peer-reviewed publications from 2017 to 2024 that directly address the intersection of philosophical theories of explanation and practical AI applications in materials discovery and property prediction. Philosophical theories of explanation, including the deductive-nomological model of Hempel and Oppenheim, the causal-mechanical account advanced by Salmon, unificationist approaches that emphasize the integration of disparate phenomena, and pragmatic frameworks that treat explanations as context-dependent answers to why-questions, provide essential benchmarks against which current materials AI practices can be assessed. In current materials AI literature, explanation is frequently conflated with prediction or post-hoc interpretability techniques such as feature importance scores and attention visualizations, as seen in comprehensive surveys of machine learning for molecular and materials science and recent advances in solid-state applications. Yet, these approaches often remain correlational rather than mechanistically grounded. XAI methods applied to materials problems, including SHAP-based feature attribution, attention mechanisms in graph neural networks, surrogate modeling, and counterfactual generation, offer valuable local insights but fall short of meeting the standards of scientific explanation due to their inherent limitations in capturing causality, multi-scale mechanisms, and physical plausibility. Ultimately, this review articulates adapted criteria for scientific explanation tailored to materials science’s multi-scale and emergent challenges and proposes actionable recommendations to bridge the gap between XAI outputs and robust explanatory accounts, urging the community to prioritize mechanistic understanding over mere predictive accuracy to advance trustworthy and insightful AI-driven discovery.
This review examines the problem of scientific consensus formation in AI-driven materials science by systematically analyzing conceptual approaches from philosophy and sociology of science alongside empirical developments in computational materials research, drawing exclusively on 31 peer-reviewed publications from 2017–2026 identified through targeted searches in Web of Science, Scopus, arXiv, and PhilPapers using terms such as “scientific consensus” AI materials, “consensus formation” machine learning science, “disagreement” materials AI, “benchmark” consensus materials informatics, “epistemic consensus” AI science, “paradigm” materials AI, “scientific disagreement” computational science, and “consensus mechanism” AI research, with inclusion criteria limited to papers addressing epistemology, disagreement, uncertainty, benchmarks, or paradigm dynamics in data-driven disciplines and exclusion of purely technical performance reports. Consensus concepts are traced from logical-positivist agreement on theories through Kuhnian paradigms and Mertonian social processes to Bayesian convergence and pragmatic problem-solving necessities, revealing how each framework illuminates different facets of knowledge coordination in materials science. AI’s impact on consensus formation operates through six distinct mechanisms—accelerated hypothesis validation, model disagreement, benchmark-driven focal points, opacity-induced dissent, data-driven convergence, and authority shifts—both facilitating rapid agreement on material properties and simultaneously generating new forms of epistemic fragmentation. These dynamics create profound tensions and paradoxes, including the trade-off between speed and deliberation, convergence versus diversity, predictive agreement versus explanatory understanding, local versus global consensus, and human versus AI authority, while exposing critical gaps such as the absence of a dedicated theory for AI-mediated consensus, the scarcity of empirical studies tracking real-time consensus processes in materials AI communities, and unresolved questions about managing productive disagreement. Recommendations are offered for researchers, journals, and the broader community to distinguish model agreement from scientific consensus, institutionalize empirical consensus studies, preserve productive dissent, and develop governance protocols that harness AI’s epistemic power without sacrificing critical scrutiny, thereby guiding the field toward more reflexive and robust knowledge production in the age of AI-augmented materials discovery.