Foundations
Section 2.3

Adversarial Features

Models may exploit predictive features that humans do not naturally use.

Adversarial examples reveal that models may exploit real predictive structure that is largely invisible or unintuitive to humans.

The conventional story

Adversarial examples are often described as arbitrary failures.

Features, not merely bugs

An alternative view is that models discover predictive features that differ from the robust features humans naturally use.

The interpretability lesson

Prediction alone does not establish that a model has learned the representation we intended.

Interpretable Intelligence synthesizes work from across the field, including research from Guide Labs. Relevant authors and results are cited throughout.