Foundations
Section 2.2

Underspecification

Why training objectives can leave important behavior unspecified.

Many learning problems admit multiple predictors with comparable in-domain performance but very different behavior under deployment conditions.

What the objective specifies

A training objective constrains only certain aspects of behavior.

What remains unspecified

Many properties we ultimately care about are not uniquely determined by the training loss.

Deployment consequences

Models that look equivalent during development can diverge substantially under distribution shift.

Interpretable Intelligence synthesizes work from across the field, including research from Guide Labs. Relevant authors and results are cited throughout.