About this collection
Interpretable Intelligence is a synthesis of ideas developed across the interpretability, machine learning, and broader AI research communities. The ideas presented here are not mine alone.
This collection is intended to distill, organize, and connect work from across the field, including—but by no means limited to—research that my colleagues and I are pursuGuide Labs.
Where possible, I will cite the original papers, authors, and results that introduced or developed the ideas discussed here. Some chapters also include my own synthesis, commentary, conjectures, and proposed frameworks; I will try to distinguish these clearly from established results.
The goal of the book is not to claim ownership over the field's ideas, but to make its accumulated knowledge—including its successes, failures, recurring mistakes, and open problems—easier to learn from and build upon.