Multi-fidelity machine learning has become a cornerstone of computational materials science because it leverages inexpensive low-fidelity data to accelerate training while reserving costly high-fidelity density functional theory (DFT) calculations for final refinement. Yet an often-overlooked source of uncertainty remains: DFT itself is noisy. Different choices of exchange-correlation functional, basis-set completeness, pseudopotential construction, and numerical convergence criteria introduce systematic and material-dependent errors that are routinely treated as exact labels. This theoretical analysis develops a unified conceptual framework for tracing how DFT error propagates through multi-fidelity training pipelines and ultimately inflates the variance of machine-learned predictions. The framework is grounded in recent theoretical and review literature on Gaussian-process and neural-network potentials, uncertainty quantification, and multi-fidelity surrogates. Proof sketches demonstrate the conditions under which multi-fidelity architectures reduce propagated variance and those under which they amplify it. Practical implications are drawn for uncertainty quantification protocols, optimal fidelity weighting, and the design of future multi-fidelity benchmarks. By making the DFT-to-ML noise pathway explicit, this work supplies a rigorous conceptual foundation for trustworthy data-driven materials modeling and highlights the necessity of reporting total (not merely model) uncertainty in high-stakes applications.
Multi-fidelity machine learning is increasingly popular in materials science: researchers routinely train surrogate models on cheap low-fidelity data (e.g., PBE-level DFT or semi-empirical tight-binding) and fine-tune on a smaller set of expensive high-fidelity DFT calculations [1-4]. The promise is clear—dramatic reduction in computational cost while retaining near-quantum accuracy [5]. Yet a hidden problem undermines this promise: DFT itself has error. Different exchange-correlation functionals (PBE versus SCAN versus hybrid or meta-GGA), incomplete basis sets, approximate pseudopotentials, and finite convergence thresholds all generate distinct energies, forces, and derived properties for the same atomic configuration. These discrepancies, often on the order of 0.1–0.5 eV per atom for formation energies and far larger for band gaps, are routinely ignored when DFT outputs are treated as ground-truth labels [6, 7].
The consequence is subtle but fundamental. A machine-learning model trained on noisy labels learns to reproduce the noisy distribution rather than the unknown true physical quantity [8]. Even a hypothetical perfect model with zero epistemic uncertainty will still output predictions whose variance is bounded below by the variance of the DFT label noise. In other words, the multi-fidelity pipeline does not eliminate error; it merely redistributes and sometimes amplifies it.
This paper provides a theoretical analysis of how noise propagates from DFT error to ML prediction variance in multi-fidelity materials models. We restrict attention to a single, transparent conceptual relationship that governs the entire process

Figure 1. Hierarchical Architecture of DFT Noise Propagation and Variance Amplification Pathways in Multi-Fidelity Materials Machine Learning
Recent literature has advanced powerful multi-fidelity architectures and sophisticated uncertainty quantification techniques [13-16], yet most treatments still assume DFT labels are exact. The present work relaxes that assumption and derives the quantitative consequences for prediction variance. By focusing exclusively on theoretical variance decomposition and proof sketches, we supply a foundational lens for assessing when multi-fidelity strategies are beneficial and when they inadvertently degrade reliability. The analysis carries direct implications for uncertainty quantification pipelines, fidelity-weighting schemes, and the future design of materials-modeling benchmarks.
Density functional theory error is more appropriately understood as a heteroscedastic, material- and property-contingent distribution rather than a single scalar quantity. A primary contribution arises from the exchange–correlation functional, which governs systematic bias: semilocal approximations such as PBE consistently underestimate band gaps by 30–100 % and formation energies by 0.1–0.5 eV/atom relative to experiment or higher-rung hybrids, while SCAN and related meta-GGAs yield improvements for certain material classes at the expense of others [6, 7]. In the absence of a universally valid functional, methodological choice itself embeds irreducible label noise. A related source of variability emerges from basis-set incompleteness, where truncation of plane-wave cutoffs, localized-basis cardinality, and k-point sampling—imposed by computational constraints—introduces residual errors on the order of 0.01–0.1 eV/atom that depend sensitively on metallicity and bonding character [17]. This dependence is further compounded by the pseudopotential or projector-augmented-wave approximation, which replaces explicit core-electron treatment with effective potentials; discrepancies across norm-conserving, ultrasoft, and PAW libraries manifest in nontrivial deviations in valence energies and forces, particularly in transition-metal and open-shell systems [7]. Beyond these approximations, self-consistent-field convergence criteria impose finite-precision limits, such that practical tolerances in energy, forces, and charge mixing leave residual fluctuations of approximately 1–10 meV/atom. Additional variance becomes pronounced in defect calculations, where supercell finite-size effects and electrostatic corrections converge slowly and remain sensitive to dielectric modeling, with correction schemes differing by 0.1–0.3 eV [14]. Under these interacting mechanisms, error magnitudes scale with both material class—more pronounced in transition-metal oxides than in covalent semiconductors—and target property, yielding intrinsically heteroscedastic noise. When such structured uncertainty propagates through multi-fidelity machine learning pipelines, it imposes a lower bound on predictive variance that cannot be reduced by model complexity alone [6, 17], making explicit characterization of DFT error indispensable for rigorous uncertainty propagation.
Table 1 decomposes DFT error into structurally distinct sources and maps each to its corresponding propagation pathway in machine-learned variance.
Table 1. Decomposition of DFT Error Sources and Their Propagation Signatures in ML Predictions
DFT Error Source | Error Type | Statistical Structure | Propagation Pathway | Impact on σ_ML² | Sensitivity to Material Class |
Exchange–correlation functional | Systematic bias | Correlated across materials | Direct inheritance + covariance | Dominant term | High (transition metals, oxides) |
Basis-set incompleteness | Numerical truncation | Weakly correlated | Additive variance | Moderate | Moderate |
Pseudopotential choice | Approximation bias | Element-specific correlation | Cross-fidelity coupling | High in alloys | High |
SCF convergence | Random residual | Weak/no correlation | Noise floor contribution | Low–moderate | Low |
Finite-size/defect corrections | Systematic + model-dependent | Structured | Amplified in sparse HF regions | High | Very high (defects) |
Property dependence | Heteroscedastic | Non-stationary | Spatial variance variation | Critical | Extreme variability |
A multi-fidelity model is defined here as any surrogate trained on data drawn from multiple sources that trade accuracy for computational cost—e.g., PBE versus SCAN DFT, or DFT versus tight-binding [18]. Noise propagation is the process by which errors in the training labels (the DFT outputs) are transferred into the statistical distribution of the model’s predictions on unseen configurations.
Consider first the single-fidelity case in which every training label originates from the same DFT level. The entire label set shares a common noise component . A machine-learning model, whether Gaussian process or neural network, is optimized to minimize loss on these noisy targets. Consequently, the learned mapping approximates the noisy DFT surface rather than the unknown true surface. Even if the model’s internal parameters are known with infinite precision (
In the multi-fidelity case the situation is richer. Low-fidelity labels (large ) are abundant and inexpensive; high-fidelity labels (smaller ) are sparse and costly [3]. The model must reconcile inconsistent targets for the same or nearby atomic environments. The optimization therefore fits a compromise surface whose variance reflects a weighted combination of the two noise sources plus their covariance. Because low-fidelity data often dominate the training set by volume, their larger error can leak into regions where high-fidelity data are absent, producing spatially heterogeneous prediction variance.
The propagation mechanism is universal: any loss function that penalizes deviation from the training labels will cause the model to partially fit the label noise. The extent of leakage is governed by the single conceptual relationship
Table 2 formalizes the distinct variance-propagation regimes that emerge from the interaction between DFT noise, model error, and cross-fidelity covariance
Table 2. Formal Regimes of Variance Propagation in Multi-Fidelity Materials Machine Learning
Regime Type | Mathematical Condition | Dominant Term | Variance Behavior | Theoretical Consequence | Practical Interpretation |
DFT-Limited | Label noise | Irreducible lower bound | Model improvement ineffective | ||
Model-Limited | σ_model² ≫ σ_DFT² | Model error | decreases with training | Classical ML regime | Data scaling beneficial |
Covariance-Amplified | Cov(HF,LF) > 0 large | Cross-term | Multi-fidelity failure | Adding LF data harmful | |
Covariance-Cancellation | Cov(HF,LF) < 0 | Cross-term | Rare beneficial regime | Requires verification | |
Weight-Misaligned | LF noise | Variance inflation | Dominance of cheap data | Reweighting required | |
Optimal Regime | Balanced weights + controlled covariance | Mixed | Theoretical optimum | Requires estimation |
Proposition 1 (Variance Decomposition for Multi-Fidelity ML). For any multi-fidelity model trained on labels that carry collective DFT noise variance , the total prediction variance obeys
Key implications follow directly. When high-fidelity DFT error dominates, further refinement of the ML architecture yields diminishing returns; the correct intervention is to adopt a higher-rung functional. When low-fidelity error is large and receives high weight, variance can increase rather than decrease. Most critically, when the covariance term is positive and large—as occurs when PBE and SCAN errors are systematically correlated across many materials—the multi-fidelity strategy can actually inflate total prediction variance compared with a high-fidelity-only baseline. Conversely, negative covariance offers a rare opportunity for error cancellation, but such cancellation is material- and property-specific and cannot be assumed a priori [6, 16 ,17].
Standard uncertainty quantification methods that report only (via ensembles, Bayesian neural networks, or Gaussian-process marginals) therefore miss the dominant contribution and systematically underestimate risk. The framework supplies an explicit lower bound on achievable prediction variance once the DFT level is fixed and clarifies why multi-fidelity performance gains reported in the literature are often smaller than expected once total uncertainty is considered [14, 15, 19].
The analysis starts from the structure of a multi-fidelity dataset, which includes high-fidelity labels formed by the unknown true value plus an error term characteristic of the high-fidelity DFT level, together with low-fidelity labels formed by the same true value plus the corresponding low-fidelity error term. The trained machine-learning model can be viewed as producing a prediction through an implicit weighted combination of these noisy labels supplemented by the model’s own approximation [20].
When the variance of this prediction is examined, it decomposes into contributions traceable to the weighted high-fidelity and low-fidelity error terms, the covariance between those error sources, and the residual model uncertainty. This decomposition aligns exactly with the framework established in Proposition 1 and rests on the core conceptual relationship
To determine whether the multi-fidelity approach yields a lower prediction variance than training exclusively on high-fidelity data, one must verify that the net contribution arising from the inclusion of low-fidelity data is smaller than the high-fidelity noise variance taken alone. Satisfaction of this condition requires that the low-fidelity error magnitude be sufficiently smaller than the high-fidelity error, that the covariance between the two error sources remain limited in its positive magnitude, and that the relative weight assigned to the low-fidelity data be selected with care.
In practice the optimal choice of these weights is governed by the precise values of the high-fidelity error variance, the low-fidelity error variance, and their covariance—quantities that are unknown a priori and must be inferred from the scarce high-fidelity data available. This estimation step is statistically challenging and represents a central practical difficulty.
The proof sketch therefore establishes two essential results. First, it demonstrates that DFT noise is inevitably inherited by any model trained on DFT-derived labels, regardless of the multi-fidelity strategy employed. Second, it identifies the narrow statistical regime in which low-fidelity data can still provide a genuine reduction in total prediction variance. The sketch also clarifies why routine application of multi-fidelity methods without explicit covariance estimation frequently produces uncertainty estimates that are poorly calibrated compared with those obtained from a high-fidelity-only baseline [16, 19].
The theoretical framework reveals four immediate consequences for uncertainty quantification in multi-fidelity materials models. First, DFT error establishes an absolute lower bound on achievable prediction variance. No matter how expressive the machine-learning architecture or how large the training set becomes, the final uncertainty cannot fall below the intrinsic noise floor set by the chosen DFT level. For applications requiring high-risk decisions—such as predicting phase stability in novel alloys or defect migration barriers in nuclear materials—this bound may render purely data-driven predictions unacceptable without additional experimental validation [14, 15, 21, 22].
Second, multi-fidelity training can paradoxically increase rather than decrease total variance. When the covariance between high-fidelity and low-fidelity errors is positive and large, the cross-term in the decomposition amplifies the propagated noise. In contrast, rare cases of negative covariance allow partial cancellation, but such cancellation is material-specific and cannot be relied upon without explicit verification. Many popular co-kriging or transfer-learning schemes implicitly assume zero covariance and therefore miss this risk [1, 3, 4].
Third, covariance is the dominant and most frequently neglected factor. Standard multi-fidelity literature often treats errors at different fidelity levels as statistically independent. In reality, errors arising from related exchange-correlation functionals (for example PBE and SCAN) are strongly correlated across broad classes of materials. Ignoring this correlation leads to over-optimistic uncertainty estimates and poor generalization [6, 16, 17].
Fourth, conventional uncertainty-quantification methods that focus exclusively on model uncertainty—ensemble variance, Bayesian neural-network marginals, or Gaussian-process predictive variances—capture only the model term and systematically underestimate the total risk [23-25].
Table 3 identifies systematic failure modes that arise when DFT noise and covariance structure are omitted from multi-fidelity analysis.
Table 3. Failure Modes in Multi-Fidelity Learning Arising from Unmodeled DFT Noise
Failure Mode | Root Cause | Mathematical Signature | Observable Symptom | Misinterpretation Risk | Required Correction |
False confidence | Ignoring | Narrow uncertainty bands | Overconfidence | Add DFT noise term | |
Variance inflation | Positive covariance | Worse than HF-only | Misattributed to model | Covariance estimation | |
Data dilution | Excess LF weighting | Performance degradation | “More data helps” fallacy | Weight optimization | |
Spatial inconsistency | Heterogeneous noise | Local prediction instability | Model blamed | Heteroscedastic modeling | |
Benchmark illusion | Treating DFT as ground truth | omitted | Inflated accuracy claims | False SOTA | Noise-aware benchmarks |
Miscalibrated UQ | Model-only uncertainty | Missing cross-term | Poor calibration | Trust in UQ misplaced | Full variance decomposition |
The DFT contribution remains invisible to these techniques, producing calibrated uncertainty only when label noise is negligible, a condition rarely met in practice [15, 19].
Taken together, these implications demand a shift in practice: uncertainty pipelines must incorporate explicit DFT-noise characterization and covariance estimation before any multi-fidelity model is deployed. Without this step, reported uncertainties remain incomplete and potentially misleading for downstream materials design.
The extent of noise propagation depends strongly on both the target property and the chemical nature of the material. Formation energies illustrate the most direct inheritance: DFT functional discrepancies of 0.1–0.5 eV/atom are transferred almost unchanged into the machine-learning prediction, setting a firm limit on the precision of phase-diagram construction or reaction-energy forecasts [6, 7].
Band-gap predictions suffer even more severely. Semilocal functionals systematically underestimate gaps by 30–100 %, an error that propagates through multi-fidelity training and contaminates any subsequent electronic-structure surrogate. The resulting prediction variance can exceed chemical accuracy for optoelectronic materials, rendering purely computational screening unreliable [17].
Elastic constants and phonon frequencies exhibit smaller DFT error (typically 5–10 %), so the propagated contribution is often tolerable for mechanical-property screening. Forces, being derivatives, generally carry lower noise than total energies; consequently, force-trained potentials inherit less DFT variance than energy-trained counterparts [26-28].
Defect formation energies represent the most challenging case. Supercell-size effects, charge corrections, and functional sensitivity combine to produce errors exceeding 0.3 eV even for well-converged calculations. In multi-fidelity settings that mix cheap low-fidelity defect data with sparse high-fidelity references, the low-fidelity noise frequently dominates distant regions of configuration space [14].
Material dependence further modulates these effects. Open-shell transition-metal compounds and oxides display the largest DFT discrepancies because of strong electron correlation and self-interaction errors. Covalent semiconductors and simple metals exhibit markedly smaller noise, allowing multi-fidelity strategies to operate closer to the ideal regime of variance reduction. This heterogeneity implies that a single uncertainty model cannot be applied universally; noise-propagation analysis must be performed property-by-property and chemistry-by-chemistry before model deployment [7, 29].
The present variance-decomposition framework complements and extends several established theoretical lines in the literature. It builds directly on the epistemic-versus-aleatoric uncertainty separation introduced in recent uncertainty-quantification studies by reframing DFT error as an irreducible aleatoric component that cannot be reduced by additional training data [15, 19]. While those works focused on model uncertainty, the current analysis quantifies how the aleatoric DFT floor propagates through hierarchical training and sets a hard limit on total predictive reliability.
Standard multi-fidelity theory, developed primarily in the context of surrogate modeling for engineering design, typically assumes that label noise is independent and identically distributed across fidelity levels. The present treatment relaxes this assumption by retaining the full covariance structure between fidelities. This relaxation reveals regimes in which classical multi-fidelity gains disappear or reverse—insights that are absent from methods that presuppose zero cross-covariance [1-4].
The framework also parallels classical error-propagation analysis in experimental physics and metrology. In physical measurements, uncertainties are combined through variance and covariance rules; machine learning adds an extra layer because the model itself can fit and amplify label noise. The single conceptual relationship therefore unifies these traditions within a data-driven materials context, showing that the ML component does not eliminate the physics-based uncertainty but merely redistributes it [14, 16].
By connecting these threads, the analysis supplies a missing conceptual bridge: it explains why reported multi-fidelity speed-ups in materials applications are often smaller than theoretical predictions once total (rather than model-only) uncertainty is considered. It also offers a clear diagnostic test—estimate the DFT covariance matrix first—to decide whether multi-fidelity is theoretically justified for a given property and material class.
For model developers the framework prescribes three concrete actions. First, characterize the DFT error distribution of every fidelity level before any training begins, using benchmark suites that span the target chemical space. Second, estimate the high-fidelity variance, low-fidelity variance, and their covariance from a small but representative set of paired calculations; these statistics then inform optimal fidelity weights that minimize propagated noise rather than merely maximizing likelihood. Third, embed the full variance decomposition into the training objective or post-hoc calibration step so that reported uncertainties reflect total rather than partial risk [16, 19].
Practitioners should adopt a more cautious workflow. Never assume that adding low-fidelity data automatically improves predictions; instead, validate every multi-fidelity model against a held-out high-fidelity test set that was never seen during weight optimization. Always report the complete uncertainty—including the DFT floor—rather than the model-only component, especially when the results will inform experimental prioritization or device design [4].
Benchmark designers must incorporate controlled DFT-error injection into future multi-fidelity challenges. Current suites treat DFT labels as ground truth; the next generation should provide both noisy and reference-level data so that propagation behavior can be quantified and compared across algorithms. Only then can the community establish which architectures truly mitigate rather than merely hide the DFT noise floor [1, 3].
Collectively these recommendations shift multi-fidelity design from an empirical art to a theoretically grounded discipline. By placing noise propagation at the center of the workflow, developers and users can avoid over-optimistic claims and build models whose uncertainty statements remain trustworthy even when the underlying DFT approximations are imperfect.
Density functional theory error propagates inevitably into the variance of machine-learned predictions whenever DFT outputs serve as training labels. The single conceptual relationship captures the entire pathway in multi-fidelity materials models, showing that even a perfect model inherits the full DFT noise floor and that the covariance term can either dampen or amplify this inheritance depending on fidelity weights and error correlations.
The analysis demonstrates that multi-fidelity strategies do not automatically reduce uncertainty; under realistic covariance conditions they can increase it. Standard uncertainty-quantification tools that ignore the DFT contribution therefore produce systematically optimistic risk assessments. Practical consequences follow directly: DFT error characterization and covariance estimation must precede any multi-fidelity deployment, optimal fidelity weights must be chosen to minimize the propagated term rather than maximize data volume, and total uncertainty—never model uncertainty alone—must be reported in all publications and applications.
By making the noise-propagation pathway explicit, this theoretical framework supplies a rigorous foundation for trustworthy data-driven materials engineering. It calls on the community to treat DFT labels as noisy observations rather than exact truths and to design the next generation of multi-fidelity algorithms with the full variance decomposition in mind. Only then can machine learning deliver the reliable, uncertainty-aware predictions that computational materials science demands.
None
None
None
None
Open Access The author(s) retain copyright. This article is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. It may be shared and adapted for non-commercial purposes with appropriate attribution, an indication of changes, and distribution of adaptations under the same license. Third-party material may be subject to separate terms identified in its credit line. View the license at https://creativecommons.org/licenses/by-nc-sa/4.0/.