Graph neural networks (GNNs) routinely achieve low errors in formation energy prediction, yet they systematically fail for metastable phases lying above the convex hull. These kinetically accessible materials, critical for batteries, catalysis, and thin-film applications, are severely underrepresented in training data dominated by stable ground states. This failure-mode analysis identifies four interlocking issues: (1) convex hull bias that pulls predictions downward toward stable energies, (2) energy range compression that collapses the predicted dynamic range, (3) local environment extrapolation failures caused by strained or unusual atomic geometries unseen in relaxed training data, and (4) error asymmetry producing consistent underprediction for metastable structures. We describe the origins, signatures, and discovery consequences of each mode, propose simple diagnostic tests for practitioners, and outline targeted mitigation strategies including metastable data augmentation, hull-aware regularization, and physics-informed extrapolation. Current GNN performance on convex-hull benchmarks does not guarantee reliability for metastable discovery. Addressing these failure modes is essential for trustworthy machine-learning-assisted design of functional materials that exist because of their position above the convex hull.