Artificial intelligence (AI) in materials science is often treated as a pipeline in which bias primarily emerges during model training, evaluation, or deployment. This framing is structurally incomplete. Many distortions later labeled as “dataset bias” are already introduced before any dataset is formally assembled, labeled, cleaned, or modeled. This conceptual manuscript advances a theory-first account of pre-dataset bias: systematic misrepresentation that originates upstream of data tables through decisions about what counts as a material instance, a property definition, a valid operating regime, and an actionable target. We argue that early bias is not merely a statistical artifact but an epistemic and procedural commitment that shapes what becomes observable, measurable, and publishable. We introduce a novel framework—the bias before data (BBD) framework—which decomposes pre-dataset bias into five coupled mechanisms: problem framing bias, regime availability bias, measurement–proxy bias, curation–visibility bias, and legitimacy bias. BBD provides a structured vocabulary for identifying where bias enters, why it persists despite technical improvements, and how it constrains the legitimacy of scientific claims even when predictive performance appears strong.