Resource hub
One curated place for what to read to get good at AI safety — including the background we deliberately don't teach. The whole point is to de-centralize the messy onboarding pipeline.
Course Material References
Texts we used & referenced in tracks.
- The case for ensuring that powerful AIs are controlledblogintermediatecontrol
Course reading in the AI Control track (Introduction, overview, and threat modeling).
- AI Control: Improving Safety Despite Intentional Subversionpaperintermediatecontrol
Course reading in the AI Control track (Introduction, overview, and threat modeling).
- Catching AIs red-handedblogintermediatecontrol
Course reading in the AI Control track (Introduction, overview, and threat modeling).
- Prioritizing threats for AI controlblogintermediatecontrol
Course reading in the AI Control track (Introduction, overview, and threat modeling).
- How can we solve diffuse threats like research sabotage with AI control?blogintermediatecontrol
Course reading in the AI Control track (Introduction, overview, and threat modeling).
- Efficient tradeoffs and the safety-usefulness tradeoff modelblogintermediatecontrol
Course reading in the AI Control track (How useful is AI control?).
- Plans A, B, C, and D for misalignment riskblogintermediatecontrol
Course reading in the AI Control track (How useful is AI control?).
- Win/continue/lose scenarios and execute/replace/audit protocolsblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- AI catastrophes and rogue deploymentsblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Rogue internal deployments via external APIsblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- A basic systems architecture for AI agents that do autonomous researchblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Practical challenges of control monitoring in frontier AI deploymentspaperintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Blocking live failures with synchronous monitorsblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Ctrl-Z: Controlling AI Agents via Resamplingpaperintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Why it's hard to make settings for high-stakes control researchblogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Untrusted advice for AI control (guided)blogintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Basic legibility protocols improve trusted monitoring (guided)paperintermediatecontrol
Course reading in the AI Control track (High-stakes control).
- Adaptive Deployment of Untrusted LLMs Reduces Distributed Threatspaperintermediatecontrol
Course reading in the AI Control track (Low-stakes control: sabotage, sandbagging, and elicitation).
- Stress-Testing Capability Elicitation (guided)paperintermediatecontrol
Course reading in the AI Control track (Low-stakes control: sabotage, sandbagging, and elicitation).
- Exploration Hacking: Can LLMs Learn to Resist RL Training?paperintermediatecontrol
Course reading in the AI Control track (Low-stakes control: sabotage, sandbagging, and elicitation).
- Sabotage Evaluations for Frontier Modelspaperintermediatecontrol
Course reading in the AI Control track (Low-stakes control: sabotage, sandbagging, and elicitation).
- Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebasespaperintermediatecontrol
Course reading in the AI Control track (Low-stakes control: sabotage, sandbagging, and elicitation).