Interpretable
Intelligence
Prologue
0.
The Priests of AGI
I · Foundations
1.
What Is Interpretability?
2.
The Underdetermination of Intelligence
3.
What Can an Explanation Tell Us?
II · Methods
4.
Input Attribution
5.
Concepts and Representations
6.
Training Data
III · Learning aControl
7.
Learning with Explanations
8.
Interpretability Efficiency and Scaling
9.
Interpretable Agents
IV · Systems
10.
Generative World Models
11.
Generative Language Models
12.
Interpretable AlphaFold
13.
Interpretable Chess and GRPO
V · Open Problems
14.
Open Problems
Learning and Control
Chapter 7
Learning with Explanations
How explanatory supervision can constrain learning and improve efficiency.
Coming soon.
← Previous
Training Data
Next →
Interpretability Efficiency and Scaling