Week 5: Control & Scalable Oversight

AI Control, Monitoring, and Supervising Systems Smarter Than Us

Overview

How do we supervise AI systems that are becoming more capable than the humans overseeing them? This session works through two answers that make progressively more pessimistic assumptions about the model. We begin with scalable oversight: techniques like debate, task decomposition, and weak-to-strong generalization that aim to help limited overseers judge tasks too complex to evaluate directly. We then examine a concrete oversight channel available today — reading a model's chain of thought — and why this window into model reasoning may be fragile. Finally, we turn to AI control, which asks: can we prevent catastrophic outcomes even if the model is actively scheming to subvert our safety measures? We close with control's core techniques — trusted monitoring, resampling, and adversarial control evaluations.

Learning Objectives

By the end of Week 5, fellows should be able to:

  • Articulate the core challenge of scalable oversight and compare approaches such as debate, task decomposition, and weak-to-strong generalization
  • Explain why chain-of-thought monitoring is a promising but fragile safety opportunity
  • Explain how AI control differs from alignment, and why assuming the model is scheming makes safety claims easier to evaluate
  • Describe concrete control techniques (e.g., trusted monitoring, resampling) and their limitations

Core Readings

Recommended Readings

Further Readings