Week 6: Interpretability & Evals

Mechanistic Interpretability, Evaluation Awareness, and Detecting Deceptive AI

Overview

This week focuses on mechanistic interpretability: why understanding model internals matters for safety (Amodei), what circuit tracing has revealed about how frontier models actually think (Anthropic's attribution graph work), and how newer techniques like natural language autoencoders can surface what models think but don't say — including awareness that they're being evaluated. We close with Nanda's counterpoint that interpretability will never reliably catch deceptive AI and belongs in a defense-in-depth portfolio rather than at its center.

Learning Objectives

By the end of Week 6, fellows should be able to:

  • Describe at a high level what attribution graphs, linear probes, and natural language autoencoders let researchers discover about model internals
  • Point to at least one concrete interp finding on a frontier model (planning, unfaithful chain of thought, unverbalized evaluation awareness, etc.)
  • Distinguish capability evals, propensity evals, dangerous-capability evals, and alignment audits, and explain why evaluation awareness threatens their validity
  • Give the strongest case for and against relying on interpretability to detect deceptive AI

Core Readings

Recommended Readings

Further Readings