Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Benchmarking Without Reality: Dataset Construction Bias in Materials Evaluation
In the rapidly evolving field of computational and data-driven materials engineering, machine learning models are increasingly deployed for property prediction, inverse design, and autonomous discovery. However, the integrity of these models hinges on the quality of training datasets, which often embed subtle biases arising from construction methodologies. This manuscript explores the conceptual underpinnings of dataset construction bias in materials AI evaluation, framing it as an epistemic challenge that distorts benchmarking outcomes and impedes genuine materials discovery. We introduce the Dataset Integrity Cascade (DIC) framework, a layered conceptual model that maps data curation processes to inference distortions, incorporating feedback mechanisms to reveal how biases propagate through representation learning, model training, and validation pipelines. By synthesizing recent advances in materials informatics, graph neural networks, and uncertainty quantification, the framework highlights systemic trade-offs between dataset scale and representational fidelity. Implications extend to high-throughput computation, closed-loop experimentation, and foundation models for science, suggesting pathways for more robust computational steering in materials design. This work underscores the need for integrative approaches that align dataset architectures with the inherent complexities of materials systems, fostering epistemically sound innovation without empirical validation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2023 | Article: 95

Failure Visibility and Epistemic Accountability in Self-Driving Materials Engineering
Self-driving laboratories have emerged as a cornerstone of computational and data-driven materials engineering, fusing automated high-throughput experimentation with machine-learning-driven decision engines to compress discovery timelines from years to weeks. This paradigm shift reconfigures the materials pipeline into a closed-loop system in which data generation, model inference, and experimental steering operate with minimal human intervention. Yet the very autonomy that accelerates discovery simultaneously obscures the epistemic foundations of the knowledge it produces. Failures—whether arising from underrepresented chemical spaces, model extrapolation beyond training distributions, or unacknowledged aleatoric–epistemic uncertainty boundaries—often remain latent until downstream validation, eroding trust in autonomous outputs. Current uncertainty quantification and explainability techniques, while technically sophisticated, are typically deployed in isolation and rarely propagate failure signals across the full discovery stack. We articulate a conceptual architecture, the Epistemic Visibility and Accountability Framework (EVAF), that treats failure not as an anomaly to be minimized but as a structured signal to be surfaced and attributed at every layer of the self-driving pipeline. By integrating multi-scale representation tracking, inference-trace logging, and risk-propagation mapping, EVAF establishes a computational substrate for epistemic accountability: the systematic assignment of responsibility for knowledge claims to specific data, model, or orchestration components. The framework reframes self-driving systems from opaque optimizers into transparent epistemic engines, enabling materials engineers to maintain intellectual oversight without sacrificing autonomy. Its implications extend to infrastructure design, regulatory readiness for autonomous discovery platforms, and the long-term reliability of data-intensive materials science.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2025 | Article: 126
Filters
Clear All





Access type