In the rapidly evolving field of materials science, artificial intelligence (AI) has emerged as a transformative tool for accelerating discovery and design. Yet, a critical bottleneck persists: the representations used to encode material properties often prioritize predictive accuracy over scientific meaning. This Perspective introduces a novel conceptual framework that bridges this gap, proposing a “Representation-Meaning Ladder” to systematically link types of representations—such as descriptors, graphs, embeddings, and text-derived variables—to the strength of scientific claims they can legitimately support, including predictive, comparative, mechanistic, causal, and transferable inferences. We argue that meaning is not inherent to representations but emerges from embedded constraints, assumptions, and contextual use, highlighting how unchecked “semantic overreach” leads to misinterpretations of correlations as mechanisms. By drawing on recent advances in materials AI, we emphasize the fragility of representations under domain shifts and the need for invariance-preserving designs to enable robust knowledge generation. This theory provides a roadmap for researchers to evaluate and enhance representations, fostering AI that not only predicts but meaningfully advances materials understanding. Ultimately, addressing this representational challenge is essential for realizing AI’s full potential in tackling complex materials challenges, from energy storage to sustainable manufacturing.
Compressed representations—such as handcrafted descriptors, autoencoder embeddings, and graph-neural-network latent spaces—have become indispensable in artificial-intelligence-driven materials science because they enable scalable property prediction from high-dimensional atomic configurations. Yet the very act of compression, while optimizing statistical correlation with target properties, systematically discards information whose scientific value lies outside mere predictive utility. This theoretical analysis applies information-theoretic principles from Shannon and Cover and Thomas to examine how dimensionality reduction in materials representations affects the retention of scientifically relevant content. Drawing on the concept of model entropy introduced by S. S., the paper introduces “model entropy” as a quantitative lens for assessing the information content preserved in any compressed materials representation. It articulates a core theoretical claim: compression optimized for predictive accuracy maximizes statistical information but can erode scientific information—mechanistic, causal, and counterfactual structures essential for understanding, explanation, and extrapolation. A typology of five distinct information-loss mechanisms is developed, each illustrated with representative materials-science scenarios. The analysis culminates in concrete implications for representation design and scientific inference, arguing that future materials AI must move beyond accuracy-centric evaluation toward explicit auditing and preservation of scientific information. By distinguishing statistical signal from epistemic content, this work offers a conceptual framework for building representations that serve both prediction and discovery without hidden epistemic costs.