Embedding domain knowledge into materials graph neural networks (GNNs) is not an optional enhancement for performance tuning; it is a fundamental requirement for models that must extrapolate reliably, operate with acceptable sample efficiency, and deliver physically consistent predictions. Purely data-driven architectures, which treat materials systems as generic graphs without built-in physical inductive biases, consistently fail when deployed beyond their training distributions — precisely the regime in which materials discovery operates. This Perspective articulates a clear position: the continued reliance on unconstrained, black-box GNNs represents a dead end for computational materials engineering. Three interlocking reasons compel this stance. First, materials discovery demands extrapolation to unseen compositions, structures, and conditions; unconstrained models offer no guarantees beyond interpolation. Second, density-functional theory (DFT) data remain scarce and computationally prohibitive, rendering sample-inefficient architectures unsustainable. Third, predictions must obey conservation laws, symmetry requirements, and thermodynamic limits; violations render long-term simulations unstable and untrustworthy. Four classes of domain knowledge must be embedded as hard constraints: (1) symmetry (E(3) equivariance, permutation invariance, space-group symmetries), (2) conservation laws (energy conservation, momentum balance), (3) physical scales and units (bounded energies, forces, and lengths), and (4) locality and smoothness principles (finite cutoffs, hierarchical many-body interactions). Recent literature provides compelling evidence that physics-constrained models — notably equivariant architectures — achieve comparable accuracy with an order-of-magnitude reduction in training data while maintaining physical consistency. Conversely, ignoring these constraints leads to extrapolation collapse, massive data waste, unphysical molecular-dynamics drift, and models that cannot be interpreted or transferred. The materials community must therefore treat physics-constrained design as the default, not an afterthought. We recommend concrete standards for model development, benchmark construction, and peer review. Only by making domain knowledge non-optional can machine learning accelerate, rather than merely decorate, the discovery of next-generation materials.
The field of artificial intelligence for materials science stands at a crossroads. After a decade of impressive leaderboard gains, it is time to confront a structural limitation that purely data-driven approaches cannot overcome.
Position Statement Embedding domain knowledge into materials GNNs is not an optional “add-on” for improving performance. It is a requirement for building models that extrapolate reliably, achieve acceptable sample efficiency, and produce physically consistent predictions. The continued pursuit of purely data-driven architectures without physics constraints is a dead end for materials discovery.
This position is not a call to abandon data or neural networks. It is a call to stop pretending that materials are simply another domain for generic graph learning. Materials data are scarce, expensive, and heavily biased toward stable, ordered structures [1-5]. Real discovery requires predicting metastable phases, disordered alloys, extreme conditions, and entirely new chemical spaces — regimes where interpolation guarantees evaporate. Pure black-box models, no matter how deep or how many parameters they contain, cannot supply the inductive biases that physics provides for free.
The urgency is immediate. High-throughput DFT campaigns have generated impressive but still limited datasets. Generating additional data for every new material family is economically and environmentally unsustainable. At the same time, industry and national laboratories demand force fields and property predictors that remain stable over nanoseconds of molecular dynamics, transfer across chemistries, and respect thermodynamic limits. Unconstrained GNNs repeatedly fail these tests. They achieve low test-set errors on random splits yet produce negative formation energies, non-conservative forces, or exploding simulations when deployed.
The community has already seen the warning signs. Early successes with message-passing networks were hailed as universal approximators, yet subsequent studies revealed that these same networks collapse outside their training distributions. The solution is not more data or larger models; it is the deliberate injection of known physics as architectural constraints [6-11]. Equivariant networks that respect E(3) symmetry, energy-conserving architectures that guarantee forces derive from a scalar potential, and locality-aware designs that mirror the finite-range nature of interatomic interactions all demonstrate superior generalization with dramatically fewer training points [12-19].
This Perspective therefore rejects the narrative that domain knowledge is merely “helpful.” It is essential. The remainder of the paper develops three reasons why constraints are non-optional, identifies the precise types of knowledge that must be embedded, marshals evidence from the recent literature, and details the concrete failures that arise when physics is ignored. The goal is not incremental improvement but a paradigm shift: physics-constrained materials GNNs must become the community standard.
Figure 1 maps the manuscript’s central argument as a directional decision structure, showing how the three core failures of unconstrained materials GNNs motivate four mandatory classes of embedded domain knowledge and culminate in concrete standards for trustworthy materials ML.

Figure 1. Why Domain Knowledge Is Non-Optional in Materials GNNs: From Structural Failures of Black-Box Models to Physics-Constrained Design Standards
The case for embedding domain knowledge rests on three interlocking failures of purely data-driven materials GNNs.
Materials discovery is intrinsically an out-of-distribution problem, as the search for novel compositions, polymorphs, and operating regimes necessarily extends beyond the support of existing datasets [10, 20]. Models trained on equilibrium structures therefore lack any principled mechanism to enforce physical validity once they depart from the training manifold. While interpolation within the convex hull of observed data may remain well behaved, predictions in extrapolative regimes often become physically implausible. The imposition of symmetry, conservation laws, and bounded physical scales constrains the hypothesis space to admissible configurations, thereby anchoring extrapolation; absent such constraints, the model remains susceptible to generating nonphysical artifacts such as negative energies or non-conservative force fields [6, 11].
This limitation becomes more acute under the stringent data regimes characteristic of atomistic modeling. Although density functional theory remains the benchmark for ground-truth data, its computational cost renders large-scale dataset expansion impractical. Within this context, the efficiency gains reported by Batzner et al. [21], where E(3)-equivariant architectures match invariant baselines using an order of magnitude fewer samples, underscore the structural advantage of embedding physical priors. Subsequent results by Batatia et al. [22] further indicate that increasing the expressivity of equivariant message passing compounds these gains. For complex material classes such as high-entropy alloys, amorphous phases, and interfacial systems, where data acquisition is severely constrained, unconstrained models effectively incur a prohibitive computational burden [18, 20]. Incorporating invariances and conservation principles thus substitutes learned regularities with encoded structure, directly reducing data requirements.
The implications extend beyond efficiency to the internal consistency of predicted physical quantities. Reliable deployment in downstream simulations requires strict adherence to foundational laws: forces must derive as gradients of scalar energies, total momentum must remain conserved in isolated systems, and predicted magnitudes must align with chemically and thermodynamically plausible scales. Data-driven models that lack these constraints frequently violate such conditions, leading to energy drift, distorted bonding geometries, and thermodynamic inconsistencies that compromise molecular dynamics trajectories. Embedding these constraints at the architectural level precludes such failures, avoiding reliance on corrective strategies applied after training.
These limitations are deeply coupled rather than independent. Deficient extrapolation reflects underlying inefficiencies in data utilization, which in turn magnify violations of physical consistency. Treating domain knowledge as optional perpetuates this feedback loop, whereas its integration restructures the learning problem itself, aligning model behavior with the governing principles of the physical systems under study.
Table 1 clarifies that the weaknesses of unconstrained materials GNNs are not isolated technical defects but structurally linked failures that directly undermine discovery-oriented deployment.
Table 1. Why Unconstrained Materials GNNs Fail on Discovery-Relevant Tasks
Discovery requirement | What the task demands | Behavior of unconstrained GNNs | Why the failure occurs | Practical consequence for materials discovery | Why embedded domain knowledge resolves it |
Extrapolation to unseen compositions, structures, and conditions | Reliable prediction outside the empirical support of training data | Performance degrades sharply once inputs move beyond equilibrium-like or previously observed regions | The model has no built-in restriction to physically admissible responses outside the training manifold | Novel chemistries, metastable phases, and extreme-condition predictions become unreliable | Symmetry, conservation, bounded scales, and locality shrink the hypothesis space to physically plausible behaviors |
Learning under DFT scarcity | Competitive accuracy with limited and expensive supervision | Requires far larger datasets to infer regularities that physics already specifies | Invariances and admissibility conditions must be learned indirectly from data rather than encoded directly | Data-generation cost becomes computationally and economically unsustainable | Hard-coded inductive biases reduce sample complexity and improve learning efficiency |
Force-field usability in simulation | Energy-force consistency and long-horizon dynamical stability | Low held-out error may coexist with unstable trajectories and drift | Outputs are not structurally tied to scalar potentials or conservation rules | Molecular dynamics becomes untrustworthy despite strong benchmark scores | Energy-conserving architectures enforce consistency by construction |
Cross-chemistry transfer | Adaptation across related but nonidentical material families | Fine-tuning is brittle and often requires extensive retraining | Learned representations are dataset-specific rather than physically grounded | Transfer learning becomes expensive, uncertain, and domain-fragmented | Physics-based constraints provide reusable structure across chemical families |
Mechanistic auditability | Ability to diagnose why the model succeeds or fails | Errors remain opaque and difficult to localize | The architecture does not separate physically meaningful sources of failure | Model debugging becomes guesswork and benchmark gains are hard to interpret | Constraint-aware designs expose violation modes such as symmetry error or conservation inconsistency |
Four classes of domain knowledge must be injected into materials GNNs as hard architectural constraints rather than soft regularizers.
Crystalline materials are governed by E(3) equivariance, wherein energies remain invariant and forces transform covariantly under rotations, translations, and inversions [16, 17, 19]. Enforcing these symmetries has been shown to substantially improve data efficiency, as demonstrated by Batzner et al. and Batatia et al. [21, 22], by aligning model structure with known physical invariances. At the same time, permutation invariance among indistinguishable atoms must be embedded directly within the message-passing mechanism, while discrete symmetries such as time-reversal in non-magnetic systems and magnetic space group constraints in spin-polarized materials further restrict the admissible function space [13]. Neglecting these conditions forces the model to expend capacity rediscovering symmetries that are already prescribed by physical law.
A related constraint emerges from conservation principles, which impose strict relationships between predicted quantities. Energy conservation requires forces to be derivable as exact gradients of a scalar potential, F_i=-∇_i E while momentum conservation enforces a vanishing net force in isolated systems, and thermodynamic consistency couples free energies, entropies, and temperature-dependent observables. Architectures that fail to encode these relations yield force fields that are fundamentally incompatible with molecular dynamics [6, 11]. By structuring the network to predict energies and derive forces analytically, such constraints are satisfied by construction rather than approximated through auxiliary objectives.
Beyond formal invariances, physically meaningful predictions must respect characteristic scales and units intrinsic to chemical systems. Bond lengths are bounded, formation energies for stable compounds lie within a narrow energetic window, and force magnitudes reflect underlying interatomic potentials. In the absence of such priors, neural networks readily produce outputs that deviate by orders of magnitude, particularly in extrapolative regimes. Incorporating scale-aware normalization, unit-consistent parameterization, and bounded representations constrains the output space and stabilizes optimization [9, 10].
These structural considerations extend naturally to the spatial organization of interactions. Interatomic forces decay rapidly beyond a few angstroms, motivating finite cutoff radii, while potential energy surfaces remain smooth except in regions associated with bond rearrangement or electronic degeneracy. Many-body interactions exhibit a hierarchical structure, with pairwise, three-body, and higher-order contributions dominating at distinct length scales. Encoding locality through cutoff-aware graph construction and reflecting this hierarchy in message passing reduces unnecessary complexity while preserving essential physics [12, 14, 15].
Taken together, these constraints delineate the admissible hypothesis space for graph neural networks in materials modeling. Omitting them does not merely degrade performance but alters the problem formulation itself, yielding models that are misaligned with the governing structure of the underlying physical system.
Table 2 consolidates the manuscript’s normative claim by specifying which forms of domain knowledge should be treated as architectural requirements, what each constraint governs, and what failure mode each prevents.
Table 2. Mandatory Classes of Domain Knowledge for Physics-Constrained Materials GNNs
Domain-knowledge class | What must be enforced architecturally | Functional role in the model | Typical implementation logic | Failure mode when omitted | Expected modeling benefit |
Symmetry constraints | E(3) equivariance, translational invariance, permutation invariance, and where relevant crystal or magnetic symmetry structure | Ensures predictions transform correctly under physically valid coordinate operations | Equivariant message passing; invariant scalar outputs for energies; vector-consistent force representations | Capacity is wasted relearning guaranteed symmetries; extrapolation becomes fragile; outputs may transform incorrectly | Stronger generalization, higher sample efficiency, and physically meaningful coordinate behavior |
Conservation laws | Forces derived from scalar energy; momentum-compatible interactions; thermodynamic consistency where relevant | Guarantees internal coherence of predictions used in downstream simulation | Energy-first architectures with analytical gradients; constraint-preserving formulations | Long-term drift, non-conservative dynamics, and simulation instability | Stable molecular dynamics and trustworthy force-field deployment |
Physical scales and units | Valid ranges, unit-consistent normalization, chemistry-aware magnitudes for energies, forces, and lengths | Prevents scale pathologies and anchors outputs to realistic material regimes | Physically informed normalization, bounded outputs, and unit-aware preprocessing or parameterization | Extreme or nonsensical outputs under distribution shift | Faster training, safer extrapolation, and more interpretable predictions |
Locality and smoothness | Finite interaction cutoffs, smooth potentials, and hierarchical many-body structure | Aligns model complexity with actual interatomic physics | Cutoff-based graph construction; local neighborhoods; layered interaction orders | Overly global or noisy representations, poor efficiency, and unstable behavior near perturbed configurations | Better inductive fit to materials physics with lower computational burden |
Constraint integration strategy | Constraints treated as hard design principles rather than optional regularizers | Determines whether physics is guaranteed or merely encouraged | Architectural enforcement preferred over penalty-only training | Constraint violations persist despite improved loss values | More robust and auditable behavior across training and deployment settings |
The literature now contains decisive evidence that physics-constrained GNNs outperform unconstrained counterparts on precisely the metrics that matter for materials discovery.
Empirical evidence consistently indicates that embedding physical structure into graph neural networks fundamentally alters their learning dynamics [5, 7, 8]. The introduction of E(3)-equivariant architectures by Batzner et al. demonstrated that NequIP attains parity with invariant models such as SchNet while requiring an order of magnitude fewer training samples, a result subsequently extended by Batatia et al. through the MACE framework, which further improves both accuracy and extrapolative robustness [21, 22]. Such gains are not attributable to incremental scaling but to the direct incorporation of symmetry constraints, a pattern replicated across crystal-graph models and higher-order equivariant formulations in diverse material settings [3, 4, 12-14, 16-19, 23, 24]. This efficiency advantage is mirrored in the behavior of energy-conserving force fields, where enforcing consistency between energies and forces suppresses the accumulation of numerical drift that otherwise leads to unphysical thermal artifacts in molecular dynamics [6, 11]. When forces are learned independently, even minor inconsistencies compound over time, whereas architectures that derive forces from scalar potentials eliminate this instability at its origin, enabling stable trajectories over extended timescales. The limitations of unconstrained approaches become particularly visible under realistic evaluation protocols: random train–test splits obscure generalization failures, as models trained on ordered crystals often collapse when applied to disordered alloys or high-temperature phases. In such regimes, the absence of physical priors prevents reliable identification of metastable configurations or extrapolation across compositional space. Although narrowly defined classification tasks, such as space-group prediction, may show limited sensitivity to these constraints, the broader literature on property prediction and interatomic potential learning leaves little ambiguity regarding their necessity.
The practical consequences of neglecting domain knowledge follow directly from these empirical patterns and manifest as a cascade of failures that compromise both reliability and scientific utility [20, 25-29]. Apparent success under interpolation masks severe degradation when models encounter novel material families, where extrapolative predictions become unreliable despite strong benchmark performance. This fragility is compounded by pronounced data inefficiency, as unconstrained architectures demand orders of magnitude more density functional theory calculations to achieve comparable accuracy, significantly inflating computational cost and extending discovery timelines. The absence of enforced physical consistency further leads to pathological behaviors, including energy drift, unstable bonding geometries, and simulations that either diverge or stagnate due to a lack of thermodynamic grounding. Interpretability deteriorates in parallel, as purely data-driven predictions offer no mechanism to disentangle errors arising from violated symmetries, broken conservation laws, or mis-scaled representations, rendering systematic debugging infeasible [25]. Transfer learning similarly degrades, with models failing to adapt across chemical domains because they lack invariant structural priors that would otherwise guide generalization. Under these conditions, unconstrained graph neural networks are not merely inefficient but fundamentally misaligned with the epistemic and operational demands of materials design.
Common objections to physics-informed modeling do not withstand scrutiny when evaluated against these constraints. Concerns that domain knowledge restricts flexibility overlook the fact that appropriate constraints, such as E(3) equivariance and energy conservation, eliminate nonphysical regions of parameter space while preserving all admissible configurations, thereby improving rather than limiting generalization, as evidenced by NequIP and MACE [21, 22]. Arguments invoking incomplete physical understanding conflate partial knowledge with absence of structure; incorporating established principles such as symmetry, locality, and conservation still yields substantial reductions in search complexity and consistently outperforms unconstrained baselines [21, 22]. Analogies to image recognition further obscure domain-specific requirements, since materials datasets are inherently sparse, costly, and extrapolation-driven, unlike the large-scale, interpolation-dominated regimes of computer vision; equivariant architectures succeed precisely because they encode this distinction [21, 22]. Implementation challenges likewise fail to justify omission, given the availability of validated libraries and the comparatively minor engineering overhead relative to the downstream cost of unstable simulations or excessive data generation [21, 22]. Finally, benchmark-centric claims of success rest on evaluation protocols that privilege interpolation and neglect physical consistency, thereby overstating model reliability. Under closer inspection, each objection ultimately reduces to a mischaracterization of the same underlying requirements—robust extrapolation, efficient data usage, and strict adherence to physical law—that render domain knowledge indispensable in materials machine learning.
This Perspective aligns with, yet sharpens, several existing strands in the literature.
We are not anti-data. Data remain indispensable for fitting the unknown residuals after physics is embedded. The position simply rejects the claim that data alone suffice. Purely data-driven GNNs treat materials as generic graphs; the result is models that scale poorly and generalize narrowly.
This is not a call to replace DFT with ML. First-principles methods will always anchor the most accurate predictions. The argument is that physics should constrain the ML layer rather than serve as its substitute. Hybrid approaches gain the speed of neural networks while inheriting the correctness guarantees of known physical laws.
The position is fully consistent with the wider machine-learning consensus that appropriate inductive biases improve generalization [26-28]. Equivariant networks are simply the materials-specific instantiation of this principle. The community has already accepted permutation invariance and cutoff locality; extending the same logic to E(3) symmetry and energy conservation is a natural, not revolutionary, step.
Physics constraints often serve as built-in interpretability mechanisms. When a model is forced to output energies from which forces are derived analytically, the causal chain becomes transparent [25]. Symmetry violations can be audited layer by layer. In contrast, black-box force predictions offer no mechanistic explanation when they fail. Embedding domain knowledge therefore advances both predictive power and scientific understanding simultaneously.
Taken together, the position does not reject any of these neighboring fields; it integrates them under a single guiding imperative: domain knowledge must be architecturally enforced, not merely encouraged.
To translate this position into community norms, we offer concrete, actionable recommendations.
Table 3 translates the Perspective’s position into an actionable evaluation standard by linking model design, benchmark construction, and review criteria to the properties that discovery-ready materials ML must actually satisfy.
Table 3. A Field Standard for Evaluating Physics-Constrained Materials GNNs
Stakeholder domain | Required evaluation question | What should be reported | Minimum acceptable evidence | What should no longer count as sufficient on its own | Intended field-level effect |
Model development | Which physical constraints are embedded, and why are they appropriate for the task? | Explicit statement of symmetry handling, conservation design, scale treatment, and locality assumptions | Architectural description plus ablation showing the effect of each constraint | Reporting predictive accuracy alone on random splits | Shifts model design from generic flexibility toward principled physical admissibility |
Benchmark construction | Does the benchmark test interpolation only, or true discovery-relevant extrapolation? | Separate results for in-distribution and out-of-distribution settings; learning curves versus data size | Novel composition, disorder, or extreme-condition test sets with transparent split logic | Single random train-test split with aggregate error metrics only | Prevents overstated generalization claims |
Physical consistency auditing | Are predictions usable in downstream simulation and decision contexts? | Force-energy consistency, symmetry-violation magnitude, long-horizon drift, and stability diagnostics | Demonstrated low violation rates and stable simulation behavior under deployment-like conditions | Low MAE without any physical-consistency reporting | Makes trustworthiness a first-class evaluation target |
Transfer and reuse | Does the model adapt across chemistry families without brute-force retraining? | Cross-domain fine-tuning or transfer results with data budgets stated explicitly | Evidence that physically grounded structure supports adaptation under limited new supervision | Claiming transferability without out-of-family tests | Rewards reusable scientific representations rather than benchmark-specific tuning |
Peer review and publishing | Can reviewers verify that the architecture is discovery-ready rather than merely leaderboard-optimized? | Constraint rationale, validation protocol, benchmark design, and code availability | Reproducible implementation and explicit answers to physics-consistency questions | Vague references to being “physics-informed” without operational evidence | Raises publication standards and improves comparability across studies |
For model developers (a) Begin every new architecture with the strongest physics constraints that the task permits—E(3) equivariance for force fields, energy conservation by construction, and locality-aware graph construction. Only after these are satisfied should additional data-driven flexibility be introduced [9, 10]. (b) Adopt equivariant message-passing layers as the default backbone for interatomic potentials [21, 22]. (c) Mandate reporting of physical-consistency metrics: force–energy consistency error, symmetry-violation magnitude, and long-term energy drift in molecular-dynamics runs.
For benchmark designers (a) Include explicit extrapolation tasks—compositionally novel alloys, disordered structures, high-pressure polymorphs—that lie outside the convex hull of training data. (b) Publish learning curves that plot error against training-set size so that sample-efficiency gains from constraints become visible. (c) Require physical-consistency tests as first-class evaluation criteria rather than optional appendices.
For journal editors and reviewers (a) Add a standard referee question: “What physics constraints are embedded in the architecture, and how were they validated?” (b) Treat purely data-driven submissions for extrapolation-critical tasks (force fields, property prediction under novel conditions) with heightened scrutiny. (c) Encourage or require open-source release of constrained reference implementations so that best practices can propagate rapidly.
Adopting these standards will shift the field from leaderboard chasing to trustworthy engineering. The cost is modest; the benefit is a generation of models that accelerate, rather than merely decorate, materials discovery.
Embedding domain knowledge into materials GNNs is not optional. It is a requirement. Purely data-driven architectures fail on extrapolation, sample efficiency, and physical consistency—the three pillars that materials discovery demands. Four classes of knowledge must be embedded as hard constraints: symmetry (E(3) equivariance and permutation invariance), conservation laws (energy and momentum), physical scales and units, and locality with hierarchical many-body interactions. Evidence from the recent literature, most notably the order-of-magnitude data-efficiency gains of NequIP and MACE, demonstrates that these constraints deliver exactly the generalization and stability that unconstrained models cannot achieve. Ignoring them produces extrapolation collapse, unsustainable data demands, unphysical simulations, opaque predictions, and poor transfer—outcomes the community can no longer tolerate.
The path forward is clear. Model developers must start with physics; benchmark designers must test extrapolation and consistency; journals must enforce these standards. Funding bodies must reward sample-efficient, physically grounded methods. Only by making domain knowledge architecturally non-optional can artificial intelligence fulfill its promise in materials science. The era of black-box GNNs for materials is over. The era of physics-constrained, trustworthy, and discovery-ready models has begun.
None
None
None
None
Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.