Institute for Advanced Materials Research Press Institute for Advanced Materials Research Press

Search

Search results:
Critical Assessment of Active Learning Benchmarks in Materials Discovery: Undisclosed Baselines and Optimistic Bias
Active learning has become a cornerstone strategy in data-driven materials discovery, promising to dramatically reduce the number of expensive simulations or experiments needed to identify high-performing materials. Proponents argue that uncertainty sampling, expected improvement, and other acquisition functions consistently outperform random selection, often by factors of 3–5× in iteration efficiency. Yet a closer examination of the benchmark studies published between 2017 and 2025 reveals a systematic pattern of optimistic bias that inflates these claims. This critical critique identifies five primary sources of overestimation: (1) weak or undisclosed random-sampling baselines, (2) unrealistic initial training sets that artificially favor active learning, (3) test-set leakage that prevents genuine extrapolation, (4) acquisition functions whose hyperparameters are implicitly tuned to the specific benchmark, and (5) incomplete reporting that hides variance and failure cases. Across the literature, random sampling is frequently presented as a naïve comparator yet proves surprisingly competitive once proper repetition, variance reporting, and realistic initial-set sizes are applied. Many studies fail to disclose the number of random seeds, the exact sampling distribution, or statistical significance tests, allowing small apparent gains to be reported as transformative. Initial training sets are often unrealistically small or already enriched with promising candidates, while test sets remain too similar to the training distribution, masking the true difficulty of exploration in vast chemical spaces. Acquisition functions are rarely subjected to hyperparameter robustness checks or evaluated on challenging out-of-distribution splits. The consequences extend beyond academic metrics: practitioners in industry and national laboratories risk deploying methods that underperform once transferred to real discovery campaigns. This critique, grounded exclusively in the 29 peer-reviewed studies listed in the reference section, calls for a new standard of rigor in active-learning evaluation. Only by adopting strong baselines, realistic initial conditions, extrapolation-aware test sets, full variance reporting, and public replication packages can the field move from optimistic benchmark theater to genuinely reliable acceleration of materials discovery.
Journal of Computational and Data-Driven Materials Engineering
Original Research | Open access | 18 July 2025 | Article: 54
Filters
Clear All





Access type