Neural networks exhibit a pronounced spectral bias, preferentially converging to low-frequency components of target functions while struggling with high-frequency details, as governed by the neural tangent kernel and the frequency principle. In the domain of materials science, accurate prediction of atomic forces in periodic crystals is essential because high-frequency force components directly encode short-wavelength phonons that govern thermal transport, lattice thermal conductivity, phonon dispersion relations, and mechanical stability under strain. Periodic crystal graphs, formed by atoms connected through cutoff-based edges under periodic boundary conditions, amplify this spectral bias relative to molecular graphs or Euclidean data: the discrete spectrum of the graph Laplacian introduces pronounced gaps between low- and high-eigenvalue modes, and standard message-passing layers act as repeated low-pass filters that exponentially attenuate high-frequency signals. We develop a theoretical framework that decomposes force fields into graph-Laplacian eigenbasis components, couples the training dynamics to the neural tangent kernel on the crystal graph, and derives bounds on high-frequency force prediction error that grow exponentially with the spectral gap of the crystal graph. This analysis reveals why conventional graph neural networks systematically miss high-frequency forces and provides rigorous mathematical mechanisms—over-smoothing, spectral-gap bottlenecks, pooling-induced filtering, and limited receptive fields—that explain the failure. The resulting bounds and propositions imply concrete architectural modifications, including residual high-frequency pathways and spectral preconditioning, that can mitigate bias while preserving the inductive biases required for periodic materials modeling.
Graph Neural Networks have become standard for predicting atomic forces in crystalline materials within the broader development of neural-network modeling [1]. However, empirical evidence shows that force prediction errors are not uniform across frequencies — high-frequency forces corresponding to short-wavelength phonons) are systematically less accurate than low-frequency forces. This paper provides a theoretical explanation for this phenomenon. We analyze the spectral bias of GNNs on periodic crystal graphs and derive bounds on high-frequency force prediction error.
In computational materials engineering, machine-learned force fields have emerged as a transformative tool for accelerating density-functional-theory-level simulations across large-scale systems. Architectures such as SchNet [2], crystal graph convolutional neural networks [3], and E(3)-equivariant graph neural networks [4] have demonstrated remarkable success in predicting energies and forces for molecules and crystals. Yet a subtle but critical limitation persists: These models exhibit markedly higher errors on high-frequency components of the force field, especially in quantum-chemistry and automated structure-manipulation contexts where fine structural variations matter [5]. High-frequency forces correspond physically to rapid spatial variations in interatomic interactions, which manifest in phonon modes with short wavelengths, large wavevectors near the Brillouin-zone boundary, and strong contributions to thermal transport and anharmonic effects.
Basri et al. showed that neural networks learn low frequencies first [6], and Xu et al. extended this analysis through the frequency principle, demonstrating that deep networks converge to band-limited approximations before capturing finer details [7]. When transferred to graph-structured data on periodic crystals, this bias is exacerbated by the intrinsic spectral properties of the crystal graph. Periodic boundary conditions discretize the Laplacian spectrum in a manner that creates large gaps between low- and high-eigenvalue subspaces, unlike the more uniform spectra encountered in finite molecular graphs or image data. Consequently, message-passing operations, which dominate modern GNN force predictors, function as low-pass filters whose repeated application erodes precisely the high-frequency information required for accurate phonon dispersion and thermal conductivity predictions.
Radovic et al. observed that machine-learned atomic forces in crystalline materials display frequency-dependent accuracy, with short-wavelength modes suffering larger deviations [8], yet provided no theoretical account of the underlying cause. Similarly, Schütt et al. demonstrated that force prediction errors correlate with vibrational frequency [2], but did not provide a theoretical explanation for this correlation. The present work closes this gap by deriving a spectral theory of GNNs on periodic crystals. We formally define spectral bias in the graph setting, link it to the eigenvalues of both the neural tangent kernel and the crystal graph Laplacian, and obtain explicit bounds showing that high-frequency force error decays at a rate bounded above by the smallest high-frequency neural tangent kernel eigenvalue, itself controlled by the Laplacian spectral gap.
This theoretical analysis proceeds in clearly delineated stages: first establishing spectral bias in generic neural networks, then specializing to periodic crystal graphs, articulating why high-frequency forces are intrinsically hard, constructing a frequency-decomposition framework coupled to neural tangent kernel dynamics, dissecting the mathematical mechanisms of suppression, and finally presenting convergence-rate bounds and design implications. Throughout, we rely exclusively on the foundational results of Basri et al. [6], Xu et al. [7, 9, 10], Li et al. [11], and the materials-specific GNN literature [2-4, 8, 12-14]. The contribution is purely theoretical: no new datasets, no empirical benchmarks, and no hyperparameter studies are introduced. Instead, we supply the missing mathematical scaffolding that explains why high-frequency forces are systematically missed and how future architectures may overcome this limitation. Table 1 maps the manuscript’s central argument across levels of analysis, showing how generic neural spectral bias is transformed into a crystal-specific failure mechanism for high-frequency force prediction.
Table 1. Multilevel Theoretical Mapping from Neural Spectral Bias to High-Frequency Force Prediction Failure in Periodic Crystals
Analytical level | Core object | Theoretical claim | Why it matters in this manuscript | Downstream implication |
Generic neural-network theory | Neural tangent kernel spectrum | Learning rates are frequency-ordered, with low-frequency components converging faster than high-frequency components | Establishes the base mechanism of spectral bias before any graph-specific structure is introduced | High-frequency target structure is intrinsically harder to fit under gradient-based training |
Graph signal representation | Graph Laplacian eigenbasis | Frequency on a crystal graph is indexed by Laplacian eigenvalues, with high-eigenvalue modes representing rapidly varying graph signals | Provides the mathematical language needed to reinterpret force fields spectrally on periodic crystal graphs | High-frequency force components become identifiable as projections onto the large-eigenvalue tail |
Periodic crystal structure | Periodic boundary conditions and discrete reciprocal-space structure | Periodicity sharpens spectral separation and creates discrete gaps between low- and high-frequency subspaces | Explains why the same spectral bias becomes more severe in crystals than in molecules or Euclidean domains | High-frequency phonon-related force information enters a structurally disadvantaged regime |
Message-passing architecture | Local aggregation layers | Standard message passing behaves as a repeated low-pass operator on graph signals | Connects architectural form directly to spectral attenuation rather than treating model failure as merely empirical | Short-wavelength force information is progressively damped across layers |
Training dynamics on crystal graphs | GNN NTK aligned with graph frequency modes | The graph-aware kernel inherits spectral ordering, so high-frequency residuals receive smaller effective learning rates | Unifies optimization dynamics with graph structure in one theoretical account | Even with continued training, high-frequency force recovery remains disproportionately slow |
Physical interpretation | Short-wavelength phonons, thermal transport, mechanical response | The most weakly learned force components are precisely those most relevant for phonon dispersion, lattice thermal conductivity, and strain-sensitive stability | Shows that the failure is scientifically consequential rather than mathematically cosmetic | Spectral bias becomes a materials-science reliability problem |
Final theoretical synthesis | Periodic-crystal GNN force prediction | Spectral failure arises from the interaction of neural bias, graph filtering, spectral gaps, and local receptive fields | Captures the manuscript’s central contribution as a composite explanation rather than a single-cause argument | Architectural redesign must explicitly preserve or reconstruct high-frequency content |
The remainder of the paper is organized as follows. Section 2 formalizes spectral bias and the frequency principle in neural networks. Section 3 defines crystal graphs and their Laplacian spectrum. Section 4 offers three informal propositions establishing the hardness of high-frequency forces. Section 5 assembles a complete theoretical framework linking Laplacian eigen-decomposition, neural tangent kernel learning dynamics, and error bounds. Section 6 details four concrete mathematical mechanisms responsible for spectral failure. (Subsequent sections will derive explicit bounds and architectural implications.) By the end, the reader will possess a rigorous, self-contained explanation of the spectral bias phenomenon in GNN-based force prediction for periodic crystals.
Figure 1 synthesizes the manuscript’s full theoretical argument by showing how generic neural-network spectral bias becomes amplified on periodic crystal graphs, generates four mechanisms of high-frequency suppression, and culminates in formal bounds and model-design implications.

Figure 1. Hierarchical Architecture of Spectral Failure in Graph Neural Networks for High-Frequency Force Prediction on Periodic Crystals
Spectral Bias — The tendency of neural networks to learn functions starting from low-frequency components and progressively moving to higher frequencies during training.
Neural Tangent Kernel — The kernel that governs the training dynamics of infinitely wide neural networks, with eigenvalues determining learning rates for different frequency components.
The frequency principle, introduced by Xu et al., provides the foundational empirical and theoretical lens for understanding spectral bias [7]. In Fourier space, the frequency principle asserts that deep neural networks fit target functions by first capturing the dominant low-frequency modes and only later, after extensive training, incorporating higher-frequency corrections. This ordering arises because the neural tangent kernel of a standard multilayer perceptron possesses eigenvalues that decay rapidly with frequency: low-frequency basis functions align with the dominant eigenvectors of the kernel, receiving exponentially faster convergence rates during gradient descent. Basri et al. rigorously quantified this phenomenon by analyzing the convergence rate of neural networks for learned functions of different frequencies and proved that the time to reach a given error tolerance scales inversely with the neural tangent kernel eigenvalue associated with each frequency band [6].
Mathematically, consider a target function defined on a compact domain and expanded in an orthonormal Fourier basis with increasing frequencies. The neural tangent kernel admits an eigen-decomposition where the eigenvalues satisfy a rapid decay for higher frequencies for typical activation functions and architectures. During training, the residual error on each frequency component decays exponentially with a rate given by the corresponding kernel eigenvalue. Consequently, low-frequency components converge exponentially faster than high-frequency components. Xu et al. later extended this analysis to show that the bias persists even in very deep networks and is reinforced by the frequency principle across a wide range of architectures [9, 10].
Cao et al. further demonstrated that the spectral bias is not merely an initialization artifact but an intrinsic property of the optimization landscape under gradient flow [15]. In the infinite-width limit, the neural tangent kernel remains constant, and training reduces to kernel regression whose solution is provably biased toward the span of the leading eigenvectors. Rahaman et al. corroborated these findings through spectral analysis on image data, showing that networks learn coarse global structure before fine local details [16].
A conceptual diagram of this phenomenon would display frequency on the horizontal axis (log scale from low to high) and learning rate on the vertical axis (log scale). The curve begins at a high plateau for low frequencies, then decays sharply after a cutoff frequency determined by network depth and width, approaching zero asymptotically for high frequencies. Superimposed would be training-time snapshots: after few epochs the low-frequency band is already near zero error, while high-frequency bands remain flat; only after orders-of-magnitude more training do the tails begin to descend. This visualization makes explicit why high-frequency targets require disproportionately long optimization and why early stopping or limited capacity exacerbates the bias.
When the target function is a force field rather than a scalar potential, the same bias applies component-wise. Because force is the negative gradient of the potential energy, high-frequency force components correspond to high-wavenumber derivatives, which inherit and amplify the spectral decay already present in the neural tangent kernel. This theoretical insight, grounded in the neural tangent kernel frequency response [6, 7], sets the stage for understanding why the bias becomes particularly severe once the input domain is replaced by the non-Euclidean geometry of a periodic crystal graph.
Crystal Graph — A graph where nodes represent atoms, edges represent interatomic bonds (typically within a cutoff radius), with periodic boundary conditions.
Graph Laplacian — the degree matrix minus the adjacency matrix, whose eigenvalues characterize the graph's frequency components.
Periodic crystals impose a discrete translational symmetry that is faithfully captured by constructing a graph with periodic boundary conditions: each atom within the unit cell is a node, and edges connect atoms whose distance is less than a cutoff radius, with images replicated across cell boundaries. The resulting graph Laplacian is circulant in the periodic directions and therefore diagonalized by discrete Fourier modes whose wavevectors lie on the reciprocal lattice. Li et al. showed that the eigenvalues of the graph Laplacian directly control the frequency content of signals propagated by graph convolutions [11]. Low eigenvalues correspond to smooth, long-wavelength modes (acoustic phonons), while high eigenvalues encode rapid oscillations (optical branches near zone boundaries).
Message-passing neural networks, the dominant paradigm for crystal force prediction, implement a local aggregation rule that combines a node's own features with a weighted sum of its neighbors' features through learned functions. In the linear case, this reduces to multiplication by a normalized adjacency matrix whose eigenvalues lie in a fixed range and are strictly smaller than one in magnitude for non-trivial graphs. Repeated application therefore damps components associated with large Laplacian eigenvalues, functioning as a low-pass filter. Schütt et al. employed precisely this message-passing construction in SchNet to achieve state-of-the-art force prediction on molecular and crystalline systems [2], while Xie et al. extended it to crystal graphs with explicit periodic encoding in crystal graph convolutional neural networks [3]. Batzner et al. later introduced E(3)-equivariant layers that preserve rotational symmetry but retain the same underlying low-pass message-passing structure [4].
The relationship between graph frequency and physical vibrational frequency is direct: the dynamical matrix for phonons is closely related to the Hessian of the potential, which itself is the second derivative of the learned energy; hence the eigenvectors of the graph Laplacian approximate the normal modes when the cutoff graph captures the relevant interaction range. High-frequency optical phonons therefore project onto high-Laplacian-eigenvalue subspaces. Because standard graph neural network layers attenuate these subspaces, the predicted force field necessarily under-represents short-wavelength vibrations.
Gilmer et al. established the message-passing framework as a general template for quantum chemistry [17], and subsequent works such as PhysNet [12] and atomistic line graph neural networks [13] confirmed that increasing depth or introducing hierarchical pooling further reinforces the low-pass character. Chen et al. analyzed simple deep graph convolutional networks and demonstrated that their effective frequency response is governed by repeated powers of the normalized Laplacian [18]. When these networks are applied to periodic crystals, the discrete translational symmetry guarantees that the Laplacian spectrum contains gaps whose size scales with the inverse square of the unit-cell size; larger cells produce denser but still gapped spectra, while small primitive cells produce large spectral gaps that separate acoustic from optical branches even more sharply.
Thus, the combination of periodic boundary conditions and message-passing induces a structural bias against high-frequency force components that is absent in non-periodic molecular graphs and molecular machine-learning benchmarks [19] or in grid-based image data where convolutional filters can be designed with arbitrary frequency response. This structural amplification of spectral bias, absent from the original neural tangent kernel analyses [6, 16], is the central reason why force prediction on crystals exhibits the observed frequency-dependent error pattern.
For a periodic crystal graph, the high-frequency components of the force field lie in the eigenspaces of the graph Laplacian with large eigenvalues.
To see this, expand the target force vector field in the orthonormal eigenbasis of the crystal graph Laplacian. Physical intuition and the dynamical matrix confirm that short-wavelength phonon modes—precisely those carrying high-frequency force information—correspond to eigenvectors with large eigenvalues near the upper end of the spectrum. Thus high-frequency forces are synonymous with projections onto the high-eigenvalue tail.
Standard message-passing graph neural networks apply a low-pass filter to node features, attenuating components corresponding to large Laplacian eigenvalues.
Consider a single linear message-passing step with the normalized adjacency operator. In the Laplacian eigenbasis, this multiplies each coefficient by a factor strictly less than one for high-frequency modes. After many layers the attenuation factor becomes very small for those modes, which for large eigenvalues decays exponentially in depth. Oono and Suzuki proved that graph neural networks exponentially lose expressive power precisely because of this eigenvalue damping [20], while Chen et al. quantified the over-smoothing effect in deep graph convolutional networks [21]. Applied to force prediction, the predicted force therefore systematically underestimates high-frequency contributions.
The neural tangent kernel of a graph neural network on a periodic crystal graph has eigenvalues that decay with graph Laplacian frequency, causing slower learning for high-frequency force components.
In the infinite-width limit, the graph neural network neural tangent kernel is a kernel on the product space of node features and graph structure. Because the message-passing operator commutes with the Laplacian eigenbasis, the neural tangent kernel inherits the same eigenbasis, and its eigenvalues decrease as the Laplacian eigenvalue increases. Xu et al. showed that the neural tangent kernel eigenvalues of graph networks decay with the graph frequency index [22], and the periodic crystal case inherits an additional discrete gap structure. Consequently, gradient descent on the high-frequency force residuals converges at a much slower rate, which is orders of magnitude slower than for low-frequency residuals.
Combining Propositions yields the central theoretical claim: high-frequency forces are theoretically hard because they live in the high-eigenvalue subspace of the crystal graph, message passing damps that subspace exponentially, and the neural tangent kernel itself assigns vanishingly small learning rates to those directions. The remainder of the paper formalizes these intuitions into a complete framework and derives quantitative bounds.
Decompose the target force field into low-frequency and high-frequency components using the graph Laplacian eigenbasis. Choose a spectral cutoff separating acoustic from optical branches; the low-frequency part collects all components below the cutoff and the high-frequency part collects all components above it.
Model the training dynamics of a graph neural network for force prediction using neural tangent kernel theory. In the infinite-width regime the parameter update follows gradient flow under the neural tangent kernel. Projecting onto the Laplacian eigenbasis diagonalizes both the loss and the kernel, decoupling the dynamics for each frequency mode.
The high-frequency error after a given number of training steps remains bounded below by an exponentially decaying term whose rate is controlled by the smallest high-frequency neural tangent kernel eigenvalue, which itself is limited by the graph Laplacian spectral gap.
Periodic boundary conditions create a discrete spectrum with gaps, making high-frequency components harder to learn than in non-periodic graphs. In a finite molecular graph the Laplacian eigenvalues are dense near zero and the spectral gap is small; in a crystal the Brillouin-zone sampling enforces larger gaps, shifting high-frequency modes farther into the damped tail of the neural tangent kernel.
Taken together, these four components constitute a closed theoretical account: the force field is spectrally decomposed, the graph neural network learns each component at a rate dictated by the neural tangent kernel eigenvalue inherited from the Laplacian, high-frequency modes receive exponentially small rates and therefore retain large residuals, and periodicity enlarges the effective spectral gap, amplifying the entire mechanism. The framework explains not only the existence of the bias but its quantitative severity in crystalline versus molecular settings.
Repeated message passing averages node features, attenuating high-frequency components exponentially with depth. In the eigenbasis, each layer multiplies coefficients by a damping factor less than one for high-frequency modes; after many layers the high-frequency amplitude decays rapidly toward zero. When predicting forces, this directly suppresses short-wavelength variations, producing overly smooth force fields. Li et al. first identified this smoothing in graph convolutional networks [11], and later analyses quantified the exponential loss of information [21, 23]. The gap between low and high eigenvalues of the graph Laplacian creates a separation in neural tangent kernel eigenvalue magnitudes. The gap induces an exponential suppression of learning rates for high-frequency modes, absent in Euclidean convolutional networks. This explains why crystal graph neural networks require far more epochs to fit optical phonons than acoustic modes.
Global or hierarchical pooling operations act as low-pass filters, discarding high-frequency information. In crystal force predictors that employ readout or hierarchical coarsening, the pooling step projects onto the constant eigenvector subspace, annihilating all components orthogonal to it. The residual high-frequency content is therefore lost irreversibly, independent of depth. Each message-passing step only aggregates local neighbors, limiting the receptive field for short-wavelength high-frequency forces. To resolve a short wavelength, the receptive field must span a sufficient distance; yet after a fixed number of layers the effective receptive field grows only linearly, while high-frequency phonons require exponentially larger fields in periodic cells. Xu et al. proved that graph neural network expressive power is bounded by the 1-WL test precisely because of this locality [24], which in the force-prediction setting translates into an inability to capture rapidly oscillating force patterns. Table 2 differentiates the four suppression mechanisms analytically, showing that the manuscript’s explanation depends on structurally distinct but mutually reinforcing sources of high-frequency loss.
Table 2. Comparative Mechanism Matrix for Spectral Suppression of High-Frequency Forces in Periodic-Crystal GNNs
Mechanism | Formal driver | Spectral action on high-frequency modes | Distinctive signature in force prediction | Why periodic crystals intensify it | Primary architectural vulnerability |
Over-smoothing | Repeated neighbor aggregation through stacked message-passing layers | Successive attenuation of large-eigenvalue components toward smooth node representations | Predicted force fields become overly smooth and under-responsive to short-wavelength variations | Crystal graphs place physically meaningful phonon information directly in the suppressed spectral tail | Increasing depth without protected high-frequency channels |
Spectral-gap bottleneck | Separation between low- and high-eigenvalue graph subspaces | Learning-rate disparity widens between low- and high-frequency residuals | Optical or boundary-zone vibrational components remain poorly fitted even when low-frequency forces converge | Discrete periodic spectra sharpen acoustic–optical separation more than in non-periodic graphs | Architectures relying entirely on inherited graph filtering and unmodified NTK ordering |
Pooling-induced low-pass filtering | Global readout or hierarchical coarsening | Projection toward coarse or smooth signal content with loss of oscillatory detail | Fine-scale force variations disappear at intermediate or output stages | Periodic crystal signals often require retention of structured local oscillation across repeated cells | Readout schemes that compress before preserving high-frequency information |
Locality of message passing | Finite receptive field growth under local aggregation | Short-wavelength interactions are not fully resolved across the spatial range needed for rapid oscillation patterns | Local force directions may appear plausible while phase-sensitive high-frequency structure is missed | Repeated unit-cell structure can require coordinated representation beyond the effective local neighborhood | Shallow local architectures without multiscale or nonlocal correction pathways |
Mechanism interaction effect | Joint action of damping, rate suppression, compression, and locality | High-frequency content is not merely learned slowly but progressively weakened in both representation and optimization | Error concentrates in the physically consequential tail rather than being uniformly distributed | Periodicity aligns physical importance with spectral disadvantage | Conventional sequential message-passing pipelines |
These mechanisms operate synergistically: over-smoothing and locality damp the signal at each layer, the spectral gap reduces the learning rate, and pooling removes whatever remains. The net effect is a systematic, theoretically unavoidable bias against high-frequency forces in any standard message-passing graph neural network applied to periodic crystals.
For a graph neural network with L layers and a message-passing operator whose eigenvalues are denoted the convergence rate for any frequency component k is proportional to raised to the power L. This means that high-frequency components, which carry eigenvalues close to the upper end of the spectrum, experience an exponentially decaying learning speed as the network depth increases. In practice, this bound shows that even moderate depth quickly renders high-frequency force components nearly unlearnable within realistic training budgets.
Under standard assumptions of infinite network width and gradient-flow training, the high-frequency force prediction error cannot fall below a positive threshold determined by the spectral gap of the crystal graph Laplacian. Specifically, once the low-frequency components have been fitted, the residual error on short-wavelength forces remains bounded away from zero by a term that scales inversely with the smallest high-frequency neural tangent kernel eigenvalue. This lower bound arises directly from the ordering of learning rates established in the frequency principle and is amplified by the discrete gaps inherent to periodic structures.
Increasing network depth improves accuracy on low-frequency forces but simultaneously degrades accuracy on high-frequency forces because of progressive over-smoothing. The trade-off is quantitative: each additional layer multiplies the high-frequency attenuation factor, producing a monotonic rise in high-frequency error even as overall training loss continues to decrease.
These bounds rest on three key assumptions. First, the network is assumed to operate in the infinite-width regime where the neural tangent kernel remains constant throughout training. Second, the crystal graph is treated as fixed with periodic boundary conditions, so the Laplacian spectrum is static and gapped. Third, the force field is expanded exactly in the Laplacian eigenbasis, an idealization that holds when the cutoff radius captures the dominant interactions. Relaxing any of these assumptions softens the bounds but does not eliminate the qualitative spectral bias. For finite-width networks, additional variance terms appear in the error, yet the leading exponential decay with frequency persists. When the cutoff radius is chosen too small, the graph becomes disconnected and the spectral gap widens further, strengthening the bounds against high-frequency learning.
The limitations of these bounds are equally important, particularly when force prediction is coupled to larger multiscale simulation settings [25]. They apply most strictly to linear or mildly nonlinear message-passing layers; highly nonlinear activations can introduce weak high-frequency pathways, but these pathways remain sub-dominant and do not overturn the overall bias. The bounds also assume mean-squared-error loss on forces; other loss functionals that explicitly weight high-frequency components could shift the effective neural tangent kernel eigenvalues but cannot remove the underlying spectral ordering. Finally, the analysis is purely theoretical and does not quantify the exact numerical value of the spectral gap for any particular material; that value must be computed from the specific crystal graph. Nevertheless, the bounds provide a universal explanation for the observed frequency-dependent errors across the entire class of periodic-crystal force predictors and establish a clear ceiling on what conventional architectures can achieve.
The spectral bias identified here connects directly to the expressive power of graph neural networks as captured by the Weisfeiler-Lehman hierarchy. Xu et al. demonstrated that standard message-passing networks are at most as powerful as the 1-WL test, which relies on iterative local aggregation and therefore cannot distinguish graphs beyond local neighborhoods. In the context of periodic crystals, this limited expressivity translates into an inability to capture short-wavelength force patterns that require global coordination across multiple unit cells. The same locality that limits graph isomorphism testing also limits high-frequency force reconstruction.
Our framework also deepens the understanding of over-smoothing and over-squashing. Li et al. first highlighted how repeated graph convolutions drive node features toward a constant vector, erasing high-frequency information. Chen et al. later quantified this smoothing through the powers of the normalized Laplacian, showing exponential decay of non-constant eigenvectors. The present analysis shows that over-smoothing is not merely a numerical artifact but the direct consequence of the neural tangent kernel inheriting the Laplacian spectrum. Over-squashing, in which distant node information is compressed into fixed-size vectors, further exacerbates the problem for high-frequency forces that oscillate rapidly across the crystal lattice.
The spectral failure also relates to graph isomorphism and symmetry considerations. Periodic boundary conditions impose additional translational symmetries that discretize the Laplacian spectrum more sharply than in finite molecular graphs. This discretization creates larger spectral gaps, which in turn widen the separation of neural tangent kernel eigenvalues. Consequently, the same symmetry that makes crystal graphs computationally convenient also makes high-frequency learning theoretically harder. Approximation theory for graph neural networks provides another link: previous results on universal approximation hold only for continuous functions on compact domains, yet they do not guarantee uniform approximation rates across frequency bands. The frequency-dependent bounds derived here supply the missing rate information and explain why approximation guarantees appear strong for energies but weaken dramatically for forces.
Taken together, these connections demonstrate that spectral bias is not an isolated phenomenon but the intersection point of several well-established theoretical limitations in graph learning. It unifies the frequency principle from classical neural networks with the spectral and topological constraints unique to graph-structured data on periodic domains.
Deeper graph neural networks worsen spectral bias. Therefore, architectures should favor shallower networks or incorporate residual connections that explicitly preserve high-frequency pathways. By adding skip connections around message-passing layers, designers can prevent the cumulative low-pass filtering that would otherwise suppress short-wavelength forces. Equivariant networks and equivariant message-passing approaches for tensorial properties [26] may partially mitigate spectral bias through better inductive biases. The E(3)-equivariant layers introduced by Batzner et al. already encode rotational symmetries that align more closely with physical force fields. Extending these layers with explicit high-frequency basis functions could further reduce the reliance on learned message passing and thereby lessen the inherited Laplacian damping. Multi-scale architectures such as graph U-nets can preserve high-frequency information through skip connections. By routing high-eigenvalue features directly from early to late layers, these models bypass the repeated attenuation that occurs in purely sequential message passing. The skip pathways act as frequency-preserving highways, allowing high-frequency force components to reach the output without being filtered out.
Fourier feature embeddings and adaptive frequency-response filtering [27] can pre-encode high-frequency components before message passing begins. By lifting node and edge features into a higher-dimensional Fourier basis that explicitly includes short-wavelength modes, the network receives a richer initial representation that counteracts the low-pass nature of subsequent layers. This preprocessing step shifts part of the spectral burden from learning to architecture design. Spectral methods that directly incorporate the graph Fourier transform may become necessary for accurate high-frequency force prediction. Instead of relying solely on local message passing, future models could perform explicit filtering in the Laplacian eigenbasis, applying learned corrections only to high-frequency modes while keeping low-frequency modes stable. Such hybrid spectral-graph networks would combine the strengths of convolutional filtering with the frequency control offered by Fourier analysis.
Table 3 converts the manuscript’s theoretical claims into an explicit design framework, showing how each identified source of spectral failure implies a distinct architectural intervention.
Table 3. Theory-to-Design Translation Framework for Mitigating Spectral Bias in Periodic-Crystal Force Predictors
Theoretical problem identified in the manuscript | Why standard GNNs fail | Design principle implied by the theory | Candidate architectural response | Expected benefit | Trade-off or caution |
High-frequency modes receive weak learning rates under NTK dynamics | Optimization favors low-frequency residuals first and disproportionately | Increase direct accessibility of high-frequency information | Residual high-frequency pathways or skip routes that bypass repeated damping | Preserves short-wavelength signal during training and inference | Can reduce simplicity and may require frequency-aware regularization |
Message passing acts as a repeated low-pass filter | Layer stacking progressively smooths large-eigenvalue components | Protect high-frequency content from cumulative attenuation | Shallow message-passing backbones with structured residual connections | Reduces over-smoothing while retaining local physical inductive bias | Too little depth may weaken low-frequency relational modeling |
Periodic crystal graphs sharpen spectral gaps | High-frequency subspaces become more isolated and harder to learn | Counteract gap-induced suppression before or during propagation | Spectral preconditioning or graph-Fourier-informed correction modules | Improves access to modes disadvantaged by graph structure | Requires additional spectral computation and graph-specific design choices |
Local aggregation cannot fully resolve rapidly oscillating patterns | Effective receptive field grows too slowly relative to the structure of short-wavelength modes | Introduce explicit multiscale or nonlocal pathways | Graph U-net variants, multiresolution pathways, or hybrid local-global architectures | Better reconstruction of oscillatory force structure across spatial scales | Added complexity may complicate training stability and interpretation |
Force targets contain physically meaningful short-wavelength content from the outset | Raw feature spaces may not expose these modes sufficiently | Encode frequency diversity before message passing begins | Fourier feature embeddings for node or edge descriptors | Makes high-frequency structure more visible to downstream layers | Poorly chosen encodings may inject irrelevant oscillatory content |
Standard architectures lack controllable frequency selectivity | The model inherits frequency response indirectly rather than designing it explicitly | Move from implicit to explicit spectral control | Hybrid spectral-graph architectures with learned correction in Laplacian space | Enables targeted improvement of the high-frequency tail without sacrificing low-frequency stability | Greater mathematical and computational burden |
These implications are not speculative; they follow rigorously from the derived bounds and mechanisms. Any architecture that fails to address the spectral gap and the low-pass character of message passing will inherit the same high-frequency limitations identified throughout this analysis [28]. The theoretical framework therefore supplies a clear roadmap for next-generation force-field models that maintain accuracy across the entire phonon spectrum.
This theoretical analysis has established that graph neural networks for periodic crystals exhibit a pronounced spectral bias: they preferentially learn low-frequency components of the target force field while systematically missing high-frequency components. The central problem arises because high-frequency forces live in the high-eigenvalue subspace of the crystal graph Laplacian, standard message-passing layers act as repeated low-pass filters, and the neural tangent kernel assigns exponentially smaller learning rates to those high-frequency directions. Periodic boundary conditions further amplify the difficulty by introducing discrete spectral gaps that separate acoustic and optical branches more sharply than in molecular graphs.
We formalized spectral bias in the graph setting, defined the crystal graph and its Laplacian spectrum, articulated three informal propositions that explain the theoretical hardness of high-frequency forces, and assembled a four-component framework that decomposes the force field, models training dynamics via the neural tangent kernel, derives error bounds, and accounts for periodicity amplification. Four mathematical mechanisms—over-smoothing, spectral-gap bottlenecks, pooling-induced filtering, and locality of message passing—were shown to operate synergistically, each contributing to the suppression of short-wavelength information. Explicit convergence-rate bounds and high-frequency error lower bounds were presented, together with a depth-frequency trade-off that quantifies the fundamental limitations of conventional architectures. These results were related to broader theoretical literature on expressive power, over-smoothing, and approximation theory, confirming that the observed bias is the natural intersection of known graph-learning constraints. Finally, five concrete implications for model design were offered, pointing toward residual connections, equivariant enhancements, multi-scale pathways, Fourier embeddings, and hybrid spectral methods as promising routes forward.
The framework and bounds supplied here close a critical gap in the theoretical understanding of machine-learned force fields. They explain, without recourse to any empirical benchmark, why high-frequency forces corresponding to short-wavelength phonons are systematically missed and why this omission matters for thermal transport, phonon dispersion, and mechanical stability. Future architecture development can now be guided by spectral considerations rather than trial-and-error alone. By incorporating explicit high-frequency preservation mechanisms, the next generation of graph neural networks for materials science can achieve uniformly accurate force predictions across all relevant frequency bands, unlocking more reliable simulations of dynamic and anharmonic phenomena in crystalline materials.
None
None
None
None
Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.