Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Bias in Materials Datasets without Datasets: A Conceptual Account of How Bias Enters Before Any Modeling
Artificial intelligence (AI) in materials science is often treated as a pipeline in which bias primarily emerges during model training, evaluation, or deployment. This framing is structurally incomplete. Many distortions later labeled as “dataset bias” are already introduced before any dataset is formally assembled, labeled, cleaned, or modeled. This conceptual manuscript advances a theory-first account of pre-dataset bias: systematic misrepresentation that originates upstream of data tables through decisions about what counts as a material instance, a property definition, a valid operating regime, and an actionable target. We argue that early bias is not merely a statistical artifact but an epistemic and procedural commitment that shapes what becomes observable, measurable, and publishable. We introduce a novel framework—the bias before data (BBD) framework—which decomposes pre-dataset bias into five coupled mechanisms: problem framing bias, regime availability bias, measurement–proxy bias, curation–visibility bias, and legitimacy bias. BBD provides a structured vocabulary for identifying where bias enters, why it persists despite technical improvements, and how it constrains the legitimacy of scientific claims even when predictive performance appears strong.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2024 | Article: 54

Conceptual Foundations for Adversarial Validation in Materials Machine Learning
Standard validation protocols in materials machine learning continue to rely on the assumption that training and test data are drawn from the same underlying distribution. This assumption is almost invariably violated in real-world materials datasets because of temporal drift in measurement techniques, compositional biases in database construction, and experimental confounders arising from different laboratories and instruments. This conceptual framework article proposes adversarial validation as a diagnostic tool specifically tailored for materials informatics: a method that trains a discriminator to explicitly detect whether a distribution shift exists between any two datasets, thereby revealing hidden generalization failures that conventional train-test splits and k-fold cross-validation cannot expose. The framework introduces the conceptual foundations of adversarial validation, distinguishes it from adversarial attacks, articulates why the technique is particularly powerful in the small-data, high-dimensional, and physically constrained domain of materials science, and offers a five-component structure for its systematic application—feature-space definition, classifier selection, shift-detection thresholding, localization of driving features, and actionable response rules. By embedding materials-specific domain knowledge into the interpretation of discriminator performance, the approach transforms validation from a passive checkpoint into an active diagnostic that can distinguish temporal shift from compositional bias and experimental confounding. The implications for materials AI practice are immediate and transformative: researchers can now report adversarial validation results alongside standard metrics, trigger targeted dataset augmentation or model retraining when shifts are detected, and document potential sources of distribution mismatch in experimental workflows, ultimately raising the robustness and trustworthiness of property predictions that underpin materials discovery and design.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 January 2022 | Article: 99

The Problem of Representational Harm in Materials Dataset Construction
Representational harm in materials dataset construction remains a critically overlooked failure mode in artificial intelligence for materials science, where systematic patterns of underrepresentation and misrepresentation silently shape which materials are studied, discovered, and deployed while rendering entire classes of materials, synthesis pathways, and knowledge traditions invisible to AI systems. Representational harm is defined here as the systematic underrepresentation, misrepresentation, or exclusion of certain material classes, synthesis methods, research traditions, or communities within materials datasets, resulting in biased AI models that perpetuate inequitable discovery outcomes and reinforce existing power structures in the field. This article articulates five distinct types of representational harm—chemical, structural, synthetic, geographic, and community—along with the four primary mechanisms through which dataset construction choices actively produce these harms, including historical priority, funding asymmetry, measurement accessibility, and publication bias. It further presents a typology of four specific harm failure modes that emerge in materials AI pipelines: invisible classes, distorted property distributions, representational feedback loops, and knowledge colonization. Finally, the paper offers practical detection principles based on diversity, geographic, citation, and community audits as well as five mitigation principles centered on intentional dataset design, data enrichment, weighted representation, inclusion of multiple knowledge systems, and ongoing harm auditing, thereby providing a comprehensive framework for transforming materials dataset construction into a more equitable and epistemically responsible practice.
Journal of Artificial Intelligence for Materials Science
Original Research | Open access | 18 July 2025 | Article: 145

Benchmarking Without Reality: Dataset Construction Bias in Materials Evaluation
In the rapidly evolving field of computational and data-driven materials engineering, machine learning models are increasingly deployed for property prediction, inverse design, and autonomous discovery. However, the integrity of these models hinges on the quality of training datasets, which often embed subtle biases arising from construction methodologies. This manuscript explores the conceptual underpinnings of dataset construction bias in materials AI evaluation, framing it as an epistemic challenge that distorts benchmarking outcomes and impedes genuine materials discovery. We introduce the Dataset Integrity Cascade (DIC) framework, a layered conceptual model that maps data curation processes to inference distortions, incorporating feedback mechanisms to reveal how biases propagate through representation learning, model training, and validation pipelines. By synthesizing recent advances in materials informatics, graph neural networks, and uncertainty quantification, the framework highlights systemic trade-offs between dataset scale and representational fidelity. Implications extend to high-throughput computation, closed-loop experimentation, and foundation models for science, suggesting pathways for more robust computational steering in materials design. This work underscores the need for integrative approaches that align dataset architectures with the inherent complexities of materials systems, fostering epistemically sound innovation without empirical validation.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 March 2023 | Article: 95
Filters
Clear All





Access type