Hands-on technical track on controlling and evaluating AI systems. The introductory module works through Redwood Research's "AI Control: Improving Safety Despite Intentional Subversion" paper and then rebuilds its trusted-monitoring result as an interactive, model-backed demo.
Sign in to track your progress and save your writing.