Why we came to accept intelligence we could not understand.
Interpretable
Intelligence
A theory and practice of understanding, communicating with, and controlling learned intelligence.
Interpretability as communication between intelligent systems.
Why successful behavior does not uniquely identify mechanism.
When explanations communicate information and when they do not.
Connecting model behavior to the inputs that influenced it.
Features, concepts, representations, and causal interventions.
Tracing model behavior back to the data from which it was learned.
How explanatory supervision can constrain learning and improve efficiency.
Measuring interpretability by cost-to-target and studying how its value scales.
Interpretation, concepts, control, and reinforcement learning in sequential agents.
Interpretable representations in generative vision and world models.
Attribution, concepts, provenance, and control in generative language models.
A canonical replication in interpretable scientific intelligence.
An end-to-end interpretable agent trained with reinforcement learning.
Unsolved problems in the science of interpretable intelligence.