Week 6: Interpretability & Evals
Mechanistic Interpretability, Evaluation Awareness, and Detecting Deceptive AI
Overview
This week focuses on mechanistic interpretability: why understanding model internals matters for safety (Amodei), what circuit tracing has revealed about how frontier models actually think (Anthropic's attribution graph work), and how newer techniques like natural language autoencoders can surface what models think but don't say — including awareness that they're being evaluated. We close with Nanda's counterpoint that interpretability will never reliably catch deceptive AI and belongs in a defense-in-depth portfolio rather than at its center.
Learning Objectives
By the end of Week 6, fellows should be able to:
- Describe at a high level what attribution graphs, linear probes, and natural language autoencoders let researchers discover about model internals
- Point to at least one concrete interp finding on a frontier model (planning, unfaithful chain of thought, unverbalized evaluation awareness, etc.)
- Distinguish capability evals, propensity evals, dangerous-capability evals, and alignment audits, and explain why evaluation awareness threatens their validity
- Give the strongest case for and against relying on interpretability to detect deceptive AI
Core Readings
Recommended Readings
- Difficulties with Evaluating a Deception Detector for AIs (GDM, 2025)
- Auditing Language Models for Hidden Objectives (Anthropic, 2025)
- A Global Workspace in Language Models (Anthropic, 2026)
- Simple Probes Can Catch Sleeper Agents (Anthropic, 2024)
- Signs of Introspection in Large Language Models (Anthropic, 2025)
- Interpretability Dreams (Olah, 2023)
- Against Almost Every Theory of Impact of Interpretability (Segerie, 2023)
- We Need a Science of Evals (Apollo Research, 2024)
Further Readings
- Toy Models of Superposition (Elhage et al., 2022)
- Towards Monosemanticity (Bricken, Templeton et al., Anthropic, 2023)
- Mapping the Mind of a Large Language Model / Scaling Monosemanticity (Templeton et al., Anthropic, 2024)
- On the Biology of a Large Language Model (Lindsey et al., Anthropic, 2025)
- Refusal in Language Models Is Mediated by a Single Direction (Arditi et al., 2024)
- Emergent Misalignment (Betley et al., ICML 2025)
- Persona Features Control Emergent Misalignment (Wang et al., OpenAI, 2025)
- AI Control: Improving Safety Despite Intentional Subversion (Greenblatt et al., ICML 2024)
- Agentic Misalignment: How LLMs Could Be Insider Threats (Lynch et al., Anthropic, 2025)
- Auditing Language Models for Hidden Objectives — full paper (Marks et al., Anthropic, 2025)
- Building and Evaluating Alignment Auditing Agents (Bricken, Marks et al., Anthropic, 2025)
- EvilGenie: A Reward Hacking Benchmark (Gabor et al., 2025)
- Sabotage Evaluations for Frontier Models (Benton et al., Anthropic, 2024)
- AI Sandbagging (van der Weij et al., 2024)
- The WMDP Benchmark (Li et al., ICML 2024)