Reading Group
Introduction
This is a small reading group working through foundational and recent papers in mechanistic interpretability; the effort to reverse-engineer the internal computations of neural networks into human-understandable algorithms. Each session centers on one paper, presented in depth, with an emphasis on building intuition for the methods (activation patching, circuit analysis, causal tracing) as well as the specific findings.
Sessions are sequenced deliberately: earlier papers introduce the tools and vocabulary (attention head taxonomy, patching, attribution) that later papers depend on, building toward more complex circuit-discovery techniques and current work on superposition and sparse autoencoders.
We expect to continue beyond the papers currently listed below – this is simply what’s been planned so far.
Useful Links
- Reading Group Calendar
- Presenter Form (fill this out if you’re interested in presenting)
- Paper Suggestion Form
- Mechanistic Interpretability Glossary
Schedule
Sessions are currently held every Wednesday from 7:30 to 8:30 p.m. Pacific Time. To register, please visit the Reading Group Calendar.
- This schedule is tentative and subject to change.
- Interactive visualizations are AI-generated and prone to errors, so please use them with caution.
Errata
If you spot any errors, please contact me at adityaiyer{dot}m{@}gmail{dot}com.