Transfer learning has become a dominant paradigm for addressing the persistent small-data limitation in computational materials science, particularly through pre-training graph neural networks on large-scale repositories and subsequent fine-tuning on task-specific datasets. This strategy assumes that universal atomic-scale representations—such as local coordination environments, bonding characteristics, and elemental embeddings—transfer effectively across domains. However, accumulating empirical evidence reveals a systematic and reproducible failure: in small-data regimes, fine-tuned models frequently underperform both zero-shot pre-trained models and models trained from scratch. This work presents a comprehensive failure-mode analysis demonstrating that such collapse is not incidental but structurally inevitable under common conditions in materials applications. Four interdependent mechanisms are identified—catastrophic forgetting, negative transfer, representation distortion, and overfitting amplification—arising from irreducible source–target domain mismatch, feature entanglement in equivariant graph architectures, and the statistical insufficiency of target datasets typically containing fewer than 10³ samples. Under these constraints, the corrective signal provided by target data is inadequate to overcome entrenched source biases or to reconfigure internal representations. Drawing exclusively on consolidated literature, we show that transfer learning failure emerges when domain similarity falls below critical thresholds and when data scarcity prevents stable optimization. The study further formalizes operational diagnostics—including baseline benchmarking, source-validation tracking, and learning-rate sensitivity analysis—and introduces mitigation strategies such as progressive unfreezing, elastic weight consolidation, replay-based regularization, and similarity-gated transfer decisions. By reframing transfer learning as a high-risk, condition-dependent strategy rather than a default solution, this work provides a principled foundation for robust model deployment in materials informatics. It challenges prevailing assumptions of universal transferability and advocates for systematic failure-mode reporting to ensure reliability in machine-learned interatomic potentials and property prediction workflows.
Domain adaptation has emerged as a promising strategy for cross-composition materials prediction, where models trained on one compositional regime are transferred to chemically distinct but related target domains. This review examines the literature from 2017 to 2024 on domain adaptation in materials informatics, with a specific focus on prediction tasks spanning binary, ternary, quaternary, and high-entropy systems. It argues that although transfer-based methods are increasingly used to reduce dependence on expensive first-principles calculations and scarce experimental measurements, their application often rests on theoretical assumptions that remain untested in materials contexts. Across the reviewed studies, six recurrent assumptions are identified: shared feature space, covariate shift only, sufficient source–target overlap, label-function stability, sample independence, and availability of target-domain data. The review shows that these assumptions are fragile under composition-induced shifts, where new bonding environments, structural prototypes, and electronic behaviors frequently alter the underlying prediction problem. It further synthesizes the field into six methodological categories—feature alignment, adversarial adaptation, reweighting, fine-tuning, multi-source adaptation, and unsupervised adaptation—and evaluates their strengths, vulnerabilities, and suitability under varying similarity regimes. Five core trade-offs structure current practice, particularly those involving performance versus negative transfer, simplicity versus flexibility, and theory versus empirical reliability. The review also identifies persistent gaps, including the lack of standardized benchmarks, weak similarity quantification, limited reporting of failure cases, and insufficient evaluation under true extrapolation. In response, it proposes a validation and reporting framework centered on pre-adaptation diagnostics, explicit similarity scoring, comparative baselines, negative-transfer reporting, and computational transparency. Overall, the review concludes that domain adaptation is most reliable as a data-efficiency tool within compositionally similar neighborhoods, but remains brittle and potentially misleading when applied across chemically distant domains without diagnostic validation.
Multi-property graph neural networks (GNNs) have emerged as the dominant paradigm in computational materials science, enabling a single equivariant backbone to simultaneously predict energies, forces, stresses, elastic tensors, band gaps, and other key properties from atomic structures. Proponents emphasize gains in data efficiency, regularization, and inductive bias transfer across tasks. However, this work reveals a fundamental and previously underappreciated limitation: equivariant error cancellation within shared representations. By enforcing rotational, translational, and permutation symmetries, equivariant GNN backbones map atomic structures into a constrained feature space where different physical properties transform under distinct irreducible representations. We demonstrate theoretically and conceptually that gradients arising from weakly correlated or anti-correlated properties frequently conflict within these shared layers. Such conflicts drive error cancellation—reductions in the aggregate multi-task loss that mask stagnant or deteriorating accuracy on individual targets—particularly when properties map to incompatible tensorial characters (scalars, vectors, rank-2 and higher tensors). We identify three regimes of representation sharing: beneficial (strongly aligned gradients), neutral (orthogonal gradients), and harmful (opposing gradients leading to representational compromise). Proof sketches show that equivariance tightens the feasible representation space and bounds gradient alignment more severely than in invariant architectures, systematically increasing the risk of harmful cancellation. Detection diagnostics and architectural mitigation strategies are proposed to identify and alleviate these effects. This analysis reframes the design of multi-property materials GNNs, demonstrating that joint training is not universally beneficial and can actively degrade specialization. It provides a principled foundation for next-generation architectures that preserve data efficiency while safeguarding per-property accuracy, advancing reliable data-driven materials discovery and design.