Open benchmark suites for machine learning interatomic potentials should include systematic failure cases as a core evaluation component rather than relying solely on leaderboard rankings based on average error metrics. Leaderboard-only benchmarks can obscure critical weaknesses, encourage overfitting to curated test sets, and provide limited guidance for real-world deployment. In contrast, failure-inclusive benchmarks reveal model limitations across challenging regimes such as compositional complexity, structural disorder, thermodynamic extremes, dynamical instability, numerical breakdown, and extrapolation beyond the training distribution. This position paper argues that making failure cases mandatory would improve transparency, support informed model selection, guide data acquisition, and strengthen reproducibility in materials machine learning. A practical typology of failure modes and a set of benchmark design principles are proposed to help shift evaluation from competitive ranking toward diagnostic understanding. Redesigning open benchmarks in this way is essential for building trustworthy and scientifically useful ML potentials.