Introduction to AI Safety DeCal — Fall 2026

The Berkeley AI Safety Initiative is running an AI Safety DeCal!

Location: Wheeler 120
Time: Mondays, 6–8 PM
First class: Monday, September 14
Units: 1

Recent incidents like the OpenAI model escaping containment and hacking HuggingFace have highlighted the real-world risks of advanced AI systems. Prominent AI researchers such as Yoshua Bengio, Stuart Russell, and Geoffrey Hinton are sounding the alarm about potential catastrophic consequences as the race to build Superintelligence accelerates. Despite this, there are not enough people focused on ensuring we develop this transformative technology responsibly.

In this DeCal, you will gain an understanding of the problems in AI safety and some of the key technical research directions that aim to solve them including Mechanistic Interpretability, Evaluations, Alignment etc. You will learn about the theoretical and practical risks associated with developing advanced AI systems, the difficulties inherent to addressing them, the current state of research regarding solutions, and various AI Safety research opportunities like Anthropic Fellows, MATS Program, SPAR, etc.

Instructors: Prakrat Agrawal, Ali Narin, Brandon Qi

Apply: Fill out the application form
Slack: Join the Berkeley AI Safety Slack

Schedule

Week Date Lecture Readings & Assignments Lecturer(s)
1 Sep 14 Logistics and Overview of AI Safety Before class In class (pick one) Reflection 1 (due before Week 2)
Full reading list
Prakrat, Ali
2 Sep 21 Machine Learning and LLM Fundamentals Before class (pick 2–3) In class (pick one) Reflection 2
Full reading list (incl. optional)
Brandon
3 Sep 28 AGI Timelines, Takeoff, and Imagining the Future Before class
  • AI 2027 (header + main timeline forecast; skim the rest)
In class (pick one) Reflection 3
Full reading list (incl. optional)
Prakrat, Ali
4 Oct 5 The Alignment Problem Before class In class (pick one) Reflection 4
Full reading list
Brandon
5 Oct 12 Inner Misalignment, Deception, and Alignment Faking Before class In class (pick one) Reflection 5
Full reading list (incl. optional)
Brandon
6 Oct 19 Control and Oversight Before class In class (pick one) Reflection 6
Full reading list (incl. optional)
Brandon, Ali (tentative)
7 Oct 26 Adversarial Robustness Before class In class (pick one) Reflection 7
Full reading list
Ali
8 Nov 2 Interpretability, Unlearning, and Representation Engineering Before class In class (pick one) Reflection 8
Full reading list (incl. optional)
Ali
9 Nov 9 Additional Research Agendas & Miscellaneous Topics Before class (pick two) In class No in-class readings.
Reflection 9 Reflection 10
Full reading list
Prakrat, Brandon
10 Nov 16 Prakrat, Brandon
11 Nov 23 Guest Speakers and Lectures on Special Topics Speakers TBD. These weeks depend on speaker availability, so guest lectures will likely be dispersed throughout the semester. Two weeks are reserved for speakers and for any delays or cancellations. TBA
12 Nov 30

This site uses Just the Docs, a documentation theme for Jekyll.