The accelerating integration of artificial intelligence into materials discovery offers transformative potential for identifying novel compounds and optimizing properties at unprecedented speeds. Yet, this promise is tempered by persistent challenges in maintaining scientific rigor across computational workflows. This review employs a structured literature synthesis grounded exclusively in 35 peer-reviewed publications from 2017 to 2025, identified through targeted searches across Web of Science, Scopus, and arXiv using strings focused on scientific rigor, reproducibility in materials machine learning, reporting standards in materials informatics, methodological quality in AI-driven science, validation standards for materials AI, benchmarking in materials property prediction, replication in computational materials science, and quality assessment frameworks for AI in materials discovery, with inclusion criteria limited to studies addressing AI-assisted discovery practices and exclusion of purely experimental or non-computational works, following a PRISMA-style screening that yielded the final corpus after removing duplicates and off-topic items. Scientific rigor in this domain is understood as the systematic application of thorough, accurate, and transparent methods that ensure independent verification of AI-generated predictions while upholding honesty in reporting both positive and negative outcomes. Current practices in materials AI demonstrate growing sophistication in model development and data utilization but reveal inconsistent transparency in code and data sharing, limited replication efforts, and reliance on internal validation that falls short of broader scientific benchmarks, even as select studies begin to engage with established checklists and principles. Critical gaps emerge in the absence of tailored materials-AI rigor frameworks, the rarity of external experimental validation, and insufficient community mechanisms for enforcing completeness in reporting, which collectively risk resource misallocation and diminished confidence in AI-driven claims. Targeted recommendations for authors, reviewers, journals, and funders emphasize mandatory code and data deposition, comprehensive hyperparameter disclosure, and cultural shifts toward valuing replication and negative results to bridge these deficiencies and elevate the field’s overall integrity.