Resources
This course pulls heavily from BlueDot Impact’s AI Alignment course, an updated version of AGI Safety Fundamentals, originally curated by OpenAI researcher Richard Ngo.
Overview / Introductory Resources
All optional — pick whatever suits your level.
- Intro to AI Safety, Remastered — 18 min video
- What happens when our computers get smarter than we are? — Nick Bostrom TED talk, 16 min
- 80,000 Hours — Preventing an AI-related catastrophe — ~60 min, quite thorough
- AGI Safety from First Principles — long document, the best technical introduction
Recommended Books
- The Alignment Problem — Brian Christian
- Superintelligence — Nick Bostrom
- Human Compatible — Stuart Russell (UC Berkeley professor)
Week 1: Logistics and Overview of AI Safety
Before class:
- What is AI Alignment? (BlueDot Impact)
- Specification Gaming: How AI Can Turn Your Wishes Against You (video)
- Why AI alignment could be hard with modern deep learning
In class (pick one):
- Agentic Misalignment: LLMs as Insider Threats
- How we could stumble into AI catastrophe
- Core Views on AI Safety: When, Why, What, and How (Anthropic)
- Reasoning through arguments against taking AI safety seriously (Bengio 2024)
Optionally, also read the other articles in the in-class list.
Week 2: Machine Learning and LLM Fundamentals
Before class (pick 2–3):
- Visualizing the Deep Learning Revolution (Ngo 2023, 20 min)
- Dive Into Deep Learning — Introduction — everything before §1.4 “Roots”; you may skip the sections on Tagging, Search, Recommender Systems, and Sequence Learning (30 min)
- 3Blue1Brown: But What Is a Neural Network? (20 min) and Gradient Descent, How Neural Networks Learn (20 min)
- Prefer text? See this intro to neural networks and this post on backpropagation
- DPO — Direct Preference Optimization
- RLHF — Reinforcement Learning from Human Feedback (InstructGPT; a variant of PPO, or Proximal Policy Optimization)
- DPO vs PPO
In class (pick one):
- GPT-3
- Understanding Reasoning LLMs
- Chinchilla scaling laws
- InstructGPT
- A Visual Guide to Mixture of Experts
Optional:
- BERT
- The Bitter Lesson
- Multi-Head Latent Attention
- QLoRA
- Challenge: train a model in PyTorch with the highest accuracy you can on MNIST (go all-out!), or work through this example notebook if you don’t know PyTorch well yet
- Supervising strong learners by amplifying weak experts (Christiano et al. 2018)
- Training a Helpful and Harmless Assistant with RLHF
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Week 3: AGI Timelines, Takeoff, and Imagining the Future
Before class:
- AI 2027 (Kokotajlo et al. 2025; read the header and main timeline forecast, then skim the other forecasts; 60 min)
In class (pick one):
- Posts from a sequence by Jacob Steinhardt: More Is Different for AI, Future ML Systems Will Be Qualitatively Different, Thought Experiments Provide a Third Anchor (20 min total)
- Situational Awareness (Aschenbrenner 2024; only read the first two sections + overview)
- Why transformative AI is really, really hard to achieve (Ramani, Wang)
- On Recent AI Model Progress
- METR: Measuring AI Ability to Complete Long Tasks (2025)
Optional:
- Dario Amodei, Machines of Loving Grace
- Eliezer Yudkowsky, The Only Way to Deal With the Threat From AI? Shut It Down
- Anthropic, Responsible Scaling Policy
Week 4: The Alignment Problem
Before class:
- Intro to AI Safety, Remastered (20 min)
- The Alignment Problem from a Deep Learning Perspective (45 min)
In class (pick one):
- MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking (Shah et al. 2025)
- The Alignment Problem from a Deep Learning Perspective
- Benefits of Assistance over Reward Learning
- Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Week 5: Inner Misalignment, Deception, and Alignment Faking
Before class:
- Deceptive alignment: ML systems will have weird failure modes (Steinhardt 2022; 10 min)
- Alignment Faking in Large Language Models (Anthropic 2024; 35 min)
- Towards evaluations-based safety cases for AI scheming (skim, 20–30 min)
In class (pick one):
- Frontier Models are Capable of In-context Scheming (Meinke et al. 2025)
- Scheming AIs: Will AIs fake alignment during training in order to get power? (Carlsmith 2023), Section 1. Carlsmith also has an online presentation of this work.
- AI Control: Improving Safety Despite Intentional Subversion (Greenblatt et al. 2024; 45 min)
- AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
- Evaluating Frontier Models for Stealth and Situational Awareness
Optional:
- Goal Misgeneralization in Deep Reinforcement Learning (Langosco et al. 2023; 30 min) — or the Rob Miles summary video (11 min)
- Parametrically Retargetable Decision-Makers Tend To Seek Power (Turner et al. 2022; 45 min)
- Risks from Learned Optimization (Hubinger et al. 2019) — or Rob Miles’ video summary, Deceptive Misaligned Mesa-Optimizers? It’s More Likely than You Think
- AI Sandbagging: Language Models can Strategically Underperform on Evaluations (van der Weij 2024)
Week 6: Control and Oversight
Before class:
In class (pick one):
- Weak-to-strong generalization (OpenAI)
- AI safety via debate
- Detecting misbehavior in frontier reasoning models (OpenAI)
- Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Activation Monitors
- AgentMisalignment: Measuring the Propensity for Misaligned Behavior in LLM-Based Agents
Optional:
- Superalignment Anti-Literature Review (Cohen et al. 2025)
- An alignment safety case sketch based on debate (Buhl et al. 2025)
- Specification gaming: the flip side of AI ingenuity (Krakovna et al. 2020; 15 min)
- Debate update: Obfuscated arguments problem (Barnes et al. 2020)
- Clearer examples of reward misspecification: The Effects of Reward Misspecification (Pan et al. 2022)
- Defining and Characterizing Reward Gaming (Skalse et al. 2022)
- One method of mitigating flawed or non-scalable oversight for RLHF: Supervising strong learners by amplifying weak experts (Christiano et al. 2018)
Week 7: Adversarial Robustness
Before class:
- Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG; Zou et al. 2023; 40 min)
- Towards Evaluating the Robustness of Neural Networks (C&W; Carlini et al. 2016; skim for 20–30 min)
In class (pick one):
- Improving Alignment and Robustness with Circuit Breakers (Zou et al. 2024)
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep (Qi et al. 2024)
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models (Liu et al. 2023)
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Trading Inference-Time Compute for Adversarial Robustness
Week 8: Interpretability, Unlearning, and Representation Engineering
Before class:
- Zoom In: An Introduction to Circuits (30 min)
- Representation Engineering (Zou et al. 2023; 45 min)
In class (pick one):
- Automatically Interpreting Millions of Features in Large Language Models
- Locating and Editing Factual Associations in GPT
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Mechanisms of Introspective Awareness
- The Platonic Representation Hypothesis
Optional:
- A Longlist of Theories of Impact for Interpretability (Nanda 2022; 7 min)
- 80,000 Hours episode #107 — Chris Olah on what the hell is going on inside neural networks (2021; 3 hours)
- Toy Models of Superposition (Elhage et al. 2022)
- Against Almost Every Theory of Impact of Interpretability
Weeks 9–10: Additional Research Agendas & Miscellaneous Topics
Before class (pick two):
- Auditing language models for hidden objectives (Anthropic 2025)
- Alignment careers guide (BlueDot Impact; Rogers-Smith 2024)
- Surfacing Pathological Behaviors in Language Models (Transluce)
- Natural Emergent Misalignment From Reward Hacking in Production RL (Anthropic)
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- Sabotage Evaluations for Frontier Models (Benton et al. 2024)
In class: no in-class readings.
Weeks 11–12: Guest Speakers and Lectures on Special Topics
These weeks depend on speaker availability, and guest lectures will thus likely be dispersed throughout the semester. We reserve two weeks of space for speakers, as well as for any possible delays or cancellations.
Other Recommended Readings
- Anthropic Alignment Science
- OpenAI Alignment Blog
- METR Research
- Request for Proposals: Technical AI Safety Research — Open Philanthropy’s overview of relevant areas and how they weigh them when funding AI safety research
- What failure looks like (Christiano 2019; 15 min)
- Clarifying “What failure looks like” (part 1) (Clarke 2020; 30 min)
- Takeoff speeds (Christiano 2018; 35 min)
- What multipolar failure looks like (Critch 2021; 45 min)
- Another outer alignment failure story (Christiano 2021)
- Value is fragile (Yudkowsky 2009; 15 min)
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition (Sharkey et al. 2025)
- hari-sikchi/awesome-ai-safety — compilation of AI safety papers
- AI Safety Papers — another compilation
Course Logistics
- Final Project guidelines — EDIT ME: add link
- Reflection submission portal — EDIT ME: add link