Machine learning (ML) has become a central driver of modern materials discovery, fundamentally reshaping how materials are designed, screened, and experimentally realized. This review examines recent advances in ML-accelerated materials discovery and emphasizes the ongoing progress in material representation and descriptor development toward fully autonomous experimental platforms. We discuss how increasingly sophisticated descriptors—ranging from composition-based features and structure-aware representations to ab initio–derived and learned embeddings—have improved predictive accuracy, data efficiency, and physical interpretability across diverse materials systems. Based on these findings, we discuss the evolution of ML frameworks for property prediction, classification, and inverse design, with particular attention to uncertainty-aware modeling, multiobjective optimization, and explainable learning strategies that bridge predictive performance with scientific insight. The study also highlights the growing role of active learning and generative models in efficiently navigating vast chemical and structural spaces, enabling data-efficient exploration and hypothesis-driven discovery. At the frontier of these developments, autonomous experimental systems integrate ML with robotics to form closed-loop workflows that iteratively design, execute, and refine experiments with minimal human intervention. Applications spanning perovskites, alloys, energy materials, and nanostructures illustrate the broad impact of these approaches in overcoming traditional trial-and-error limitations. Finally, we discuss persistent challenges associated with data scarcity, extrapolation, interpretability, and system integration, and outline future directions toward more robust, scalable, and sustainable autonomous materials discovery. Collectively, these advances represent a paradigm shift from passive data-driven prediction to intelligent, self-guided materials innovation.
In the rapidly evolving field of materials science, artificial intelligence (AI) has emerged as a transformative tool for accelerating discovery and design. Yet, a critical bottleneck persists: the representations used to encode material properties often prioritize predictive accuracy over scientific meaning. This Perspective introduces a novel conceptual framework that bridges this gap, proposing a “Representation-Meaning Ladder” to systematically link types of representations—such as descriptors, graphs, embeddings, and text-derived variables—to the strength of scientific claims they can legitimately support, including predictive, comparative, mechanistic, causal, and transferable inferences. We argue that meaning is not inherent to representations but emerges from embedded constraints, assumptions, and contextual use, highlighting how unchecked “semantic overreach” leads to misinterpretations of correlations as mechanisms. By drawing on recent advances in materials AI, we emphasize the fragility of representations under domain shifts and the need for invariance-preserving designs to enable robust knowledge generation. This theory provides a roadmap for researchers to evaluate and enhance representations, fostering AI that not only predicts but meaningfully advances materials understanding. Ultimately, addressing this representational challenge is essential for realizing AI’s full potential in tackling complex materials challenges, from energy storage to sustainable manufacturing.