Generative materials models, including variational autoencoders, generative adversarial networks, and diffusion models, have become central to modern artificial intelligence for materials science. Yet, their pervasive reliance on analogy-based reasoning remains largely unexamined and conceptually undertheorized. These models routinely treat latent-space interpolation, transfer learning, and structural substitution as forms of analogical mapping—assuming that what holds between known materials will hold for novel ones—without acknowledging the fundamental epistemological limits of such reasoning. This critical critique identifies four interlocking problems that undermine the reliability of analogy-driven generation: analogy functioning as a substitute for genuine physical understanding, the propagation of false analogies, boundary blindness to domains where analogies break, and the reification of statistical correlations into ontological claims. The consequences of these unacknowledged limits extend beyond technical inaccuracy to wasted experimental resources, overconfident predictions, and a subtle distortion of scientific understanding in materials discovery. Rather than abandoning analogy entirely, this paper argues for hybrid frameworks that explicitly bind analogical transfer with physical invariants, causal verification, and uncertainty quantification. By confronting these conceptual limits head-on, the field can move toward more robust, epistemologically grounded generative models that augment rather than replace mechanistic insight.
The advent of foundation models, large-scale pre-trained architectures adapted from natural language processing paradigms, has permeated computational materials science, promising accelerated discovery through data-driven inference. In materials engineering, these models leverage multimodal datasets encompassing atomic structures, properties, and simulations to enable representation learning across scales. However, inherent conceptual limits arise from the interplay between materials' physical hierarchies—spanning quantum to macroscopic levels—and the inductive biases embedded in pretraining strategies. This manuscript synthesizes recent advancements in machine learning architectures, such as graph neural networks and multimodal integration, within materials informatics ecosystems. It identifies epistemic boundaries where foundation models falter in capturing causality, uncertainty, and domain-specific invariances, potentially leading to misaligned discovery pipelines. To address these, we introduce the Matter Pretraining Boundary Framework (MPBF), a conceptual architecture that delineates layers of data assimilation, representational abstraction, and inference steering to mitigate limits in autonomous materials design. Implications extend to high-throughput computation, inverse design, and simulation-experiment coupling, fostering more robust computational workflows in materials engineering. By interpreting these limits through systems-level dynamics, the framework guides infrastructure trade-offs, enhancing the reliability of data-driven paradigms without empirical validation.