Interpretable
Intelligence
Prologue
0.
The Priests of AGI
I · Foundations
1.
What Is Interpretability?
2.
The Underdetermination of Intelligence
3.
What Can an Explanation Tell Us?
II · Methods
4.
Input Attribution
5.
Concepts and Representations
6.
Training Data
III · Learning aControl
7.
Learning with Explanations
8.
Interpretability Efficiency and Scaling
9.
Interpretable Agents
IV · Systems
10.
Generative World Models
11.
Generative Language Models
12.
Interpretable AlphaFold
13.
Interpretable Chess and GRPO
V · Open Problems
14.
Open Problems
Methods
Chapter 6
Training Data
Tracing model behavior back to the data from which it was learned.
Coming soon.
← Previous
Concepts and Representations
Next →
Learning with Explanations