Equivariant graph neural networks have become central to data-efficient modeling of interatomic potentials, yet the mechanism underlying their empirical advantage remains insufficiently quantified. This work establishes a theoretical link between E(3)/SE(3) equivariance and sample complexity reduction in materials GNNs. By constraining the hypothesis space to functions that respect physical symmetries, equivariant architectures eliminate orbit-related redundancy and reduce the effective dimension of the learning problem. Within a PAC-style framework, we derive a bound showing that the sample complexity of an equivariant hypothesis class is reduced relative to its unrestricted counterpart by a factor proportional to the symmetry group size |G|, up to logarithmic terms. For discrete crystallographic groups, this yields finite reductions (e.g., 48× for cubic symmetry), while for continuous groups such as SO(3), the reduction becomes unbounded in the finite-data regime. The analysis relies on group-averaging projections and the induced contraction of VC dimension. These results provide a quantitative explanation for the observed data efficiency of architectures such as NequIP, MACE, and Allegro, and clarify the limitations of non-equivariant baselines such as SchNet. The framework situates symmetry as a statistically optimal inductive bias and offers principled guidance for designing data-efficient models in materials discovery.
Equivariant graph neural networks (GNNs) have rapidly become the preferred architecture for learning interatomic potentials and electronic properties in crystals, molecules, and amorphous materials [1-3]. This shift reflects a broader trend in scientific machine learning: the move from flexible but unconstrained models toward architectures that explicitly encode known physical structure [4]. In atomistic systems, one of the most fundamental sources of such structure is symmetry. Physical observables are not arbitrary functions of atomic coordinates; they obey strict transformation rules under translations, rotations, and reflections. Models that fail to respect these symmetries must either learn them from data—often inefficiently—or risk producing physically inconsistent predictions.
Architectures that incorporate E(3) or SE(3) equivariance—most notably NequIP [5] and MACE [6]—embed these symmetry constraints directly into their design. As a result, they ensure that predictions transform correctly under changes of coordinate frame, eliminating entire classes of unphysical solutions by construction. Empirically, this design choice has led to dramatic improvements in both accuracy and data efficiency. For instance, NequIP [5] trained on roughly 100 structures can match or exceed the performance of SchNet [7] trained on an order of magnitude more data. GemNet [8] and related directional message-passing networks [9, 10] exhibit similar trends, achieving competitive or superior performance while relying on significantly fewer training examples. These results are now consistently observed across molecular benchmarks, bulk materials datasets, and increasingly complex chemical environments.
While this empirical success is well documented, its underlying cause is often described only qualitatively—typically attributed to “better inductive bias” or “incorporation of symmetry.” Such explanations, while intuitively appealing, do not quantify how much is gained by enforcing equivariance, nor do they clarify the precise mechanism by which data efficiency improves. This gap between empirical observation and theoretical understanding motivates the present work.
At a conceptual level, the key effect of equivariance is to constrain the hypothesis space of the learning problem. In supervised learning, the hypothesis space consists of all functions that a model can represent given its architecture and parameterization. For standard (non-equivariant) GNNs, this space is extremely large and contains many functions that violate the symmetry constraints of physical systems. Consequently, the learning algorithm must use data to both identify the target function and implicitly infer the underlying symmetries. This dual burden increases the number of training examples required and can lead to inefficient generalization, particularly in low-data regimes.
In contrast, equivariant models restrict the hypothesis space to functions that are consistent with the action of a symmetry group G. This restriction has two important consequences. First, it removes redundant degrees of freedom: functions that differ only by symmetry transformations are no longer treated as distinct. Second, it enforces a form of automatic data augmentation: learning from a single configuration implicitly provides information about all configurations related by the group action. Together, these effects reduce the effective complexity of the learning problem.
From the perspective of statistical learning theory, this reduction in complexity translates directly into lower sample complexity. Sample complexity is defined as the minimal number of labeled examples needed to ensure, with high probability, that the learned model achieves a target generalization error ε. Classical results show that this quantity scales with the effective dimension of the hypothesis class, as measured by quantities such as the VC dimension. By constraining the hypothesis space through equivariance, the effective dimension is reduced, leading to a corresponding decrease in the number of samples required for reliable learning.
Figure 1 presents the manuscript’s central theoretical logic as a directional hierarchy: physical symmetry constrains the admissible hypothesis class, collapses orbit-related redundancy, contracts effective dimension, and thereby reduces sample complexity.

Figure 1. Hierarchical mechanism of symmetry-induced sample complexity reduction in equivariant materials GNNs
The central contribution of this work is to make this intuition precise. We derive a PAC-style bound showing that the sample complexity of an equivariant hypothesis class H_eq is reduced relative to that of the unrestricted class H_full by a factor proportional to the size of the symmetry group |G| (up to logarithmic corrections). This result provides a concrete and interpretable scaling law: the larger the symmetry group, the greater the reduction in required training data. For example, in cubic crystal systems with 48 symmetry operations, the bound predicts a 48-fold reduction in sample complexity. In systems with continuous rotational symmetry, the effective reduction is even more substantial, reflecting the infinite structure of the group and the corresponding collapse of redundant representations.
This theoretical result offers a direct explanation for the empirical performance of modern equivariant architectures. Models such as NequIP [5], MACE [6], and Allegro [11] succeed in low-data regimes because they operate within a dramatically smaller and more physically meaningful hypothesis space. In contrast, non-equivariant models such as SchNet [7] must approximate symmetry relationships through exposure to large and diverse datasets, effectively “wasting” samples on learning transformations that could otherwise be enforced by design [12]. Directional models like GemNet [8] partially address this issue by incorporating geometric information [13], but still lack full equivariance, and therefore do not achieve the same degree of sample efficiency.
Beyond explaining existing architectures, our analysis builds on and extends prior theoretical work on symmetry in machine learning [14]. Early contributions by Keriven and Peyré [15] established the general framework of group-equivariant neural networks, while subsequent studies examined the statistical benefits of invariance in more abstract settings [16, 17]. However, these results do not directly translate to the specific case of atomistic modeling, where continuous symmetries, geometric constraints, and graph-based representations interact in nontrivial ways [18]. By deriving explicit bounds tailored to materials GNNs, the present work bridges this gap and connects abstract group-theoretic principles with practical model design.
In summary, this work shows that equivariance provides not only a physically meaningful inductive bias but also a quantifiable statistical advantage. By aligning the hypothesis space with the symmetry structure of physical systems, equivariant GNNs achieve superior generalization with significantly fewer training examples. This insight has important implications for computational materials science, where high-quality labeled data is often expensive to obtain [19], and suggests a general strategy for designing data-efficient models in other domains governed by symmetry.
A function f is said to be equivariant with respect to a group G if, for every group element g ∈ G and input x, the following holds:
Invariance is the special case in which the output remains unchanged:
Exact E(3)-equivariant GNNs such as NequIP [5], MACE [6], and Allegro [11] embed this property directly into the architecture via irreducible representations of the rotation group and equivariant message passing. In contrast, SchNet [7] and early GemNet variants [8] are only approximately invariant: they rely on the training data to learn directional information, which is statistically costly. Tensor field networks and E(n) GNNs [20] provide the foundational continuous-group framework that later materials-specific models refined.
Conceptually, consider a cubic crystal structure. Apply any of the 48 operations of the O_h group. An equivariant model produces force vectors that transform exactly as the input coordinates. An approximately invariant model, however, may output inconsistent directions unless it has seen enough rotated copies during training.
Higher-order equivariant networks such as those in [21-24] further demonstrate that increasing the order of tensor representations preserves equivariance while expanding expressivity within the symmetry-constrained subspace. The key statistical consequence is that the model no longer treats rotated or reflected configurations as independent examples; they are identified by construction. This identification is the mechanism by which sample complexity is reduced.
Sample complexity refers to the smallest number of training examples required such that, with high probability (at least 1−δ), a learned hypothesis generalizes to within a target error tolerance ϵ [25]. In other words, it quantifies how much data is needed before a model can be trusted to make accurate predictions on unseen inputs. Within the Probably Approximately Correct (PAC) framework, this quantity is not determined solely by the learning algorithm, but depends fundamentally on the structure and richness of the hypothesis class HHH from which the model is selected.
A central insight from statistical learning theory is that richer hypothesis classes—those capable of representing a wider variety of functions—require more data to learn reliably. This is because a larger function class increases the risk of overfitting: with too much flexibility, the model can fit noise or spurious patterns in the training data rather than capturing the true underlying relationship. Measures such as the VC dimension formalize this notion of capacity by quantifying how complex a hypothesis class is in terms of the patterns it can represent. As this effective dimension increases, so too does the number of samples required to constrain the model and ensure good generalization.
It follows directly that any mechanism which reduces the effective complexity of the hypothesis class will also reduce the required number of training examples. Equivariant architectures provide exactly such a mechanism. By enforcing symmetry constraints, they restrict the hypothesis class from the full space of arbitrary functions to a smaller subset consisting only of symmetry-respecting functions, denoted Heq. This restriction is not merely a heuristic regularization; it systematically eliminates entire families of functions that are inconsistent with the underlying physics. As a result, the learning problem becomes intrinsically simpler, since the model no longer needs to consider or rule out invalid hypotheses during training.
The connection to materials GNNs is immediate and particularly striking. The space of all possible message-passing functions defined on atomic graphs is extremely large, encompassing many functions that do not respect rotational or translational symmetry. A non-equivariant model operating in this space must effectively use data to discover the correct transformation behavior, learning, for example, that rotating a molecular configuration should not change its energy. This requirement places a significant burden on the training process, as the model must observe many symmetry-related configurations before it can internalize these invariances.
Equivariant GNNs eliminate this burden by embedding the symmetry transformation rules directly into the architecture. Instead of learning symmetry from data, the model is constructed so that its outputs automatically transform correctly under the action of the symmetry group. This architectural constraint dramatically shrinks the search space of possible functions, reducing redundancy and enabling the model to generalize from far fewer examples. In effect, each observed training configuration carries more information, because its implications extend across all symmetry-related configurations without requiring explicit data augmentation.
This dimension-reduction perspective provides a unifying explanation for the empirical success of modern equivariant models. Architectures such as NequIP [5], MACE [6], and related approaches [11, 26, 27] consistently demonstrate strong performance in low-data regimes precisely because they operate within a constrained and physically meaningful hypothesis space. By contrast, models that do not enforce equivariance must compensate with larger datasets, longer training times, or additional regularization strategies [23]. Viewed through this lens, the advantage of equivariant GNNs is not incidental, but a direct and predictable consequence of reduced hypothesis class complexity.
Without built-in symmetry, a model must learn that every rotated or translated copy of a crystal produces correspondingly transformed outputs. Each such copy lies on the orbit of the original structure under the group G. For a discrete group the orbit size is exactly |G|; for SO(3) the orbit is a continuous manifold whose volume is the Haar measure of the group.
A non-equivariant hypothesis class treats each orbit member as an independent input–output pair. Consequently, the effective number of distinct degrees of freedom scales with |G| × (number of unique structures). An equivariant class, by contrast, identifies all orbit members automatically. The hypothesis space therefore collapses by a factor of |G| (or the group volume).
Formally, let denote the dimension (or VC dimension) of the unrestricted class . The equivariant subclass satisfies
For a cubic crystal (|G| = 48 under ) the reduction is 48-fold. The model effectively “sees” 48 times more data because symmetry-augmented copies are no longer distinct examples that must be memorized. This orbit-stabilizer argument appears in various forms in the invariant-learning literature [16, 17] and is here specialized to materials GNNs.
The reduction is even stronger for continuous groups. The space of non-equivariant functions must approximate an infinite-dimensional orbit; without equivariance the learner cannot generalize rotational symmetry from any finite dataset. Equivariant architectures circumvent this impossibility theorem by construction.
Let be any hypothesis class of functions on atomic configurations and let
Corollary For crystals possessing the full cubic symmetry group (|G| = 48), an exactly equivariant GNN requires roughly 48× fewer training structures than a non-equivariant counterpart to reach identical generalization error. For SE(3)-equivariant models the reduction factor diverges in the finite-data limit, explaining why non-equivariant networks cannot reliably learn force fields without exhaustive data augmentation.
The bound holds because the VC dimension contracts exactly as derived in Section 4. It directly accounts for the superior data efficiency reported for NequIP [5], MACE [6], Allegro [11], and related E(3)-equivariant architectures [20, 21, 26, 27].
The proof proceeds in five transparent steps using only group averaging and standard VC-dimension arguments.
Step 1. Any candidate function f can be projected onto the equivariant subspace by the group-average operator:
Step 2. Each orbit of size |G| contributes only a single independent degree of freedom to , because the values on all orbit members are rigidly linked by the equivariance condition. Consequently the VC dimension satisfies
Step 3. Substitute the reduced dimension into the PAC bound:
Step 4. Replacing d by immediately yields
Step 5. For continuous groups the averaging integral over SO(3) replaces the finite sum; the orbit volume becomes infinite, rendering formally infinite while remains finite. Hence non-equivariant models require infinite data in the worst case.
The argument extends results on invariant networks [16, 17] and sample-complexity bounds for equivariant architectures [15, 28] to the graph-structured, material-specific setting. Limitations include the worst-case nature of VC bounds and the assumption of uniform group averaging; distribution-dependent improvements may be even larger.
Non-equivariant models such as SchNet [7] and early variants of CGCNN rely on the optimizer to discover approximate rotational invariance from data. They must allocate capacity and samples to learn that rotated inputs produce rotated outputs—an expensive process that inflates the effective hypothesis-space dimension. Equivariant models (NequIP [5], MACE [6], Allegro [11]) embed the transformation rule directly, freeing capacity for learning physically relevant variations.
Table 1 distinguishes the statistical consequences of three materials-GNN hypothesis classes and clarifies why exact equivariance yields the sharpest contraction in effective dimension and sample complexity.
Table 1. Theoretical comparison of unrestricted, approximately invariant, and exactly equivariant materials GNN hypothesis classes
Dimension of comparison | Unrestricted / non-equivariant GNNs | Approximately invariant / directional GNNs | Exactly E(3)- or SE(3)-equivariant GNNs |
Representative model logic | Generic message passing without enforced transformation law | Geometric features improve directional sensitivity, but symmetry is not guaranteed exactly | Transformation law embedded directly in architecture |
Relation to physical symmetry | Symmetry must be inferred from data | Partial or approximate accommodation of symmetry | Exact respect for rotations, translations, and relevant reflections by construction |
Status of orbit-related configurations | Treated as distinct training cases | Partially linked through learned geometric patterns | Identified automatically as symmetry-related instances |
Hypothesis-space structure | Largest admissible function class | Intermediate restriction of function class | Smallest physically admissible function class |
Redundant degrees of freedom | High | Reduced but still substantial | Strongly suppressed through symmetry constraints |
Need to learn rotational behavior from samples | High | Moderate | Minimal, because transformation behavior is built in |
Effective dimension | dfulld_{\text{full}}dfull | Between deqd_{\text{eq}}deq and dfulld_{\text{full}}dfull | (d_{\text{eq}} \approx d_{\text{full}}/ |
Expected sample complexity scaling | Highest | Intermediate | Lowest; theoretically reduced by approximately (1/ |
Role of augmentation | Often necessary to approximate symmetry | Still beneficial in many settings | Largely unnecessary for exact symmetry enforcement |
Risk of physically inconsistent outputs | Highest | Reduced but not eliminated | Lowest within the symmetry-constrained target class |
Best-fit regime | Large-data settings tolerant of statistical inefficiency | Medium-data settings with partial geometric structure | Data-scarce scientific learning with strong known symmetry |
Empirical validation cited from the literature confirms the predicted gap: NequIP achieves force MAE of 0.02 eV/Å with only 100 training structures, whereas SchNet requires approximately 5000 structures to reach comparable accuracy [1, 4]. The theoretical bound supplies the missing quantitative link, showing that the observed gap is not merely architectural but follows directly from the |G|-fold contraction of sample complexity.
The bound derived in Proposition 1 sits naturally within the broader landscape of statistical learning theory while providing a symmetry-specific sharpening absent from classical results. The standard PAC bound based on VC dimension, as recalled in Section 3, gives the generic scaling m ≈ (d/ϵ) log(1/δ) without regard to structure in the hypothesis class. The present analysis exploits the group-theoretic structure of to replace with
The result also aligns with Rademacher-complexity analyses of symmetric functions. Bietti et al. [16] showed that invariance reduces Rademacher complexity by a factor related to orbit size, yet their bound applies primarily to discrete permutation groups and does not address the continuous E(3) case central to materials GNNs. Elesedy and Zaidi [17] established a provably strict generalization benefit for equivariant models under certain margin conditions; the present work recovers their qualitative improvement as a special case while quantifying the precise factor |G| via VC-dimension contraction rather than margin arguments.
Information-theoretic perspectives on group averaging further support the bound. The operation
A related but distinct line appears in invariant risk minimization (IRM), where the goal is to learn representations invariant across environments. While IRM seeks robustness to distribution shifts, the present bound is purely structural: it holds even under a single fixed data distribution because the symmetry constraint is baked into the architecture. The key distinction is quantification: earlier bounds state that “symmetry helps,” whereas Proposition 1 supplies the concrete factor 1/|G| (or ∞ for continuous groups) that explains the dramatic data-efficiency gap observed between NequIP [5], MACE [6], and non-equivariant baselines such as SchNet [7].
Collectively, the bound unifies and sharpens these threads by specializing group-theoretic dimension reduction to the E(3)-equivariant message-passing setting of materials GNNs. It thereby bridges abstract statistical-learning theory with the concrete engineering constraints of data-scarce computational materials discovery.
The theoretical reduction factor |G| carries immediate, actionable consequences for the architecture and deployment of materials GNNs. First, equivariance is not an optional architectural flourish but a statistical necessity when training data are limited—a regime that dominates high-throughput materials screening. Model developers should therefore prioritize E(3)- or SE(3)-equivariant message passing (as realized in NequIP [5], MACE [6], and Allegro [11]) over approximately invariant baselines whenever the target property transforms as a vector, tensor, or higher-order irreducible representation. Partial symmetry (for example, rotation invariance without full vector equivariance) yields only a subgroup reduction factor and should be viewed as a fallback rather than a target.
Second, for practitioners facing datasets smaller than 1000 structures, the bound supplies a quantitative rule of thumb: select an exactly equivariant architecture to obtain the full 48-fold (cubic) or effectively infinite (continuous) reduction in required samples. Non-equivariant models such as SchNet [7] or early GemNet variants [8] will require data-augmentation pipelines whose computational cost scales with |G|, negating much of their apparent simplicity. In contrast, built-in equivariance renders explicit augmentation unnecessary, simultaneously lowering training time and memory footprint.
Third, benchmark designers must shift from final-accuracy tables to sample-efficiency curves (error versus training-set size). Reporting only converged performance on large datasets obscures the very advantage that Proposition 1 predicts. Future benchmarks should include logarithmic-scale learning curves and explicitly compute the empirical reduction factor at fixed error thresholds, thereby validating the theoretical |G| scaling across crystal systems [29].
Finally, the analysis highlights an opportunity for hybrid designs: when full E(3) equivariance is computationally prohibitive for very large systems, one may embed lower-order equivariant layers early in the network and recover higher-order tensors only where needed. The bound guarantees that even partial symmetry injection still contracts sample complexity by the corresponding subgroup order, providing a principled trade-off between expressivity and data efficiency. In short, the theory reframes equivariance from a “nice-to-have” physical constraint into a provably optimal inductive bias for the data-scarce world of materials discovery.
This work establishes a direct connection between symmetry and sample efficiency in materials graph neural networks. By restricting the hypothesis space to functions that respect E(3) or SE(3) equivariance, the learning problem eliminates orbit-related redundancy and reduces its effective complexity. This structural constraint explains why equivariant architectures achieve strong generalization with substantially fewer training examples.
The analysis provides a theoretical foundation for the observed data efficiency of modern equivariant models while clarifying the limitations of non-equivariant approaches. Rather than treating equivariance as a heuristic inductive bias, the results show that it induces a principled contraction of the hypothesis space that directly improves statistical efficiency.
The implications are methodological. In data-constrained regimes typical of computational materials science, equivariance should be treated as a default architectural constraint rather than an optional design choice. More broadly, the results demonstrate how embedding physical symmetry into model design yields predictable gains in generalization, offering a general strategy for developing data-efficient machine learning systems in symmetry-governed scientific domains.
None
None
None
None
Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.