Many learning problems admit multiple predictors with comparable in-domain performance but very different behavior under deployment conditions.
What the objective specifies
A training objective constrains only certain aspects of behavior.
What remains unspecified
Many properties we ultimately care about are not uniquely determined by the training loss.
Deployment consequences
Models that look equivalent during development can diverge substantially under distribution shift.