Resources

This course pulls heavily from BlueDot Impact’s AI Alignment course, an updated version of AGI Safety Fundamentals, originally curated by OpenAI researcher Richard Ngo.

Overview / Introductory Resources

All optional — pick whatever suits your level.


Week 1: Logistics and Overview of AI Safety

Before class:

In class (pick one):

Optionally, also read the other articles in the in-class list.

Week 2: Machine Learning and LLM Fundamentals

Before class (pick 2–3):

In class (pick one):

Optional:

Week 3: AGI Timelines, Takeoff, and Imagining the Future

Before class:

  • AI 2027 (Kokotajlo et al. 2025; read the header and main timeline forecast, then skim the other forecasts; 60 min)

In class (pick one):

Optional:

Week 4: The Alignment Problem

Before class:

In class (pick one):

Week 5: Inner Misalignment, Deception, and Alignment Faking

Before class:

In class (pick one):

Optional:

Week 6: Control and Oversight

Before class:

In class (pick one):

Optional:

Week 7: Adversarial Robustness

Before class:

In class (pick one):

Week 8: Interpretability, Unlearning, and Representation Engineering

Before class:

In class (pick one):

Optional:

Weeks 9–10: Additional Research Agendas & Miscellaneous Topics

Before class (pick two):

In class: no in-class readings.

Weeks 11–12: Guest Speakers and Lectures on Special Topics

These weeks depend on speaker availability, and guest lectures will thus likely be dispersed throughout the semester. We reserve two weeks of space for speakers, as well as for any possible delays or cancellations.


Course Logistics


This site uses Just the Docs, a documentation theme for Jekyll.