The integration of physical principles into machine learning (ML) frameworks has emerged as a transformative approach in materials science, addressing the limitations of purely data-driven models by incorporating domain knowledge to enhance predictive accuracy, generalizability, and interpretability. This narrative review explores the conceptual taxonomies of physics-integrated ML methods, their applications in materials discovery and design, and the associated challenges in data bias and ethical considerations. Drawing on recent peer-reviewed literature, we classify physics-integration strategies such as physics-informed neural networks (PINNs), hybrid models combining ML with physical simulations, and constraint-based learning, and highlight their roles in solving complex problems such as material property prediction, microstructure analysis, and phase stability. We also examine how data biases in training datasets can propagate errors and inequities in model outputs, and discuss the ethical values underpinning the use of AI in scientific research, including transparency, accountability, and societal impact. The review underscores the potential of these methods to accelerate innovation in materials science while emphasizing the need for rigorous validation and interdisciplinary collaboration. By synthesizing current advancements, this article aims to provide a foundational understanding for researchers and practitioners, paving the way for future developments in this interdisciplinary field.
Ensemble methods have become the dominant framework for uncertainty quantification in machine learning models for high-entropy alloy (HEA) property prediction, where variance across independently trained neural networks is routinely interpreted as epistemic uncertainty. This metric now underpins active learning, compositional screening, and experimental decision-making, largely due to its simplicity and success in data-rich domains. This work shows that such reliance is fundamentally misplaced in HEAs. Ensemble variance implicitly assumes IID sampling, feature-space isotropy, uniform error, independence among models, and Gaussian residuals—conditions that are systematically violated in compositionally complex alloys. HEA datasets are biased toward equiatomic, stable compositions, the compositional manifold is anisotropic, predictive error is strongly heteroscedastic, ensemble members exhibit correlated failures, and extrapolation induces heavy-tailed errors. Under these conditions, ensemble variance becomes miscalibrated, underestimating uncertainty in sparse regions while overstating model reliability. The resulting distortions propagate through discovery workflows, yielding inefficient active-learning strategies, overconfident extrapolation, and misleading experimental guidance, even as larger ensembles appear to reduce uncertainty without improving accuracy. By linking these failures to their statistical origins, this paper clarifies why standard ensemble diversity is not a faithful proxy for epistemic uncertainty in HEAs. It further outlines diagnostics and targeted corrections that align uncertainty estimates with the physics and data structure of complex alloy systems, enabling more reliable and efficient materials discovery.
Machine learning has transformed materials engineering through graph neural networks, equivariant architectures, generative models, and early autonomous laboratories. These advances enabled accurate property prediction, data-efficient force fields, and initial inverse design of inorganic crystals, supported by community databases such as the Materials Project and JARVIS. Generative frameworks now propose novel structures, while closed-loop platforms demonstrate early integration of prediction, synthesis, and robotic experimentation. However, major gaps remain: weak extrapolation beyond training distributions, inadequate handling of long-range interactions, poorly calibrated uncertainty quantification, limited synthesis prediction, underrepresentation of disordered materials, and insufficient experimental validation. Persistent unlearned lessons include benchmark biases that inflate generalization performance, poor reproducibility practices, misuse of invariant models for tensor properties, omission of random baselines in active learning, weak validity filtering in generative workflows, and near-absent failure reporting. This review critically assesses a decade of progress and shortcomings, grounded in 38 peer-reviewed publications. It highlights that hype around foundation models, generative inverse design, and autonomous labs often exceeds demonstrated impact. Realizing the field’s potential requires prioritizing extrapolation-focused architectures, mandatory experimental validation, shared failure registries, and stronger evidentiary standards to accelerate reliable discovery of functional materials.