M8: AI Safety and Monitoring
As autonomous systems grow more capable, monitoring and oversight become critical. This module covers interpretability, alignment, and the submodular optimization methods that deploy monitoring resources efficiently. Building on M3 trigger strategies and M6 evaluation methods, we study how to detect deviations from expected behavior and design oversight that scales. PA5 makes you the AI controller: you will build a monitoring system for a multi-agent environment.
Module preview
As autonomous systems grow more capable, monitoring and oversight become critical. This module covers interpretability, alignment, and the submodular optimization methods that deploy monitoring resources efficiently. Building on M3 trigger strategies and M6 evaluation methods, we study how to detect deviations from expected behavior and design oversight that scales. PA5 makes you the AI controller: you will build a monitoring system for a multi-agent environment.
Lectures and materials
L17: AI Safety Monitoring
Interpretability, alignment, and detecting unexpected behavior.
🔜 Fall slides forthcomingL18: Submodular Optimization for Monitoring
Efficient allocation of limited oversight resources.
🔜 Fall slides forthcoming
Programming assignment
PA5: AI Control · Due December 5 at 11:59 PM CDT