Bridging Theory, Understanding, and Practice
Fall 2026 | UC Berkeley
Large-Scale AI as an End-to-End Engineering Discipline
Central inquiry: How do we build, train, and deploy large-scale AI systems by treating them as full-stack engineered artifacts, where hardware constraints, software stacks, and optimization dynamics jointly determine model behavior and performance?
This course examines the principles required to build, train, and deploy large-scale AI models. We treat large-scale AI as an end-to-end engineering discipline, where a model is a computational graph that must be trained, specialized, evaluated, deployed, monitored, and iterated on, under hard constraints from hardware, data, and serving economics.
The course follows the lifecycle end to end:
NVIDIA is the compute sponsor for Scalable AI: Bridging Theory, Understanding, and Practice. Their support provides the GPU infrastructure that makes the course labs and projects possible.
Ed only. Please post all course questions on Ed (course discussion board). We do not use Slack or other channels for course Q&A. If you have not been automatically added, the join code will be shared in lecture.
Once we finalize the enrollment, we will purge all non-students from the course forum to encourage free discussion among students.
Teams are 8 groups of 4. Each group receives 1 H100 node (8ΓH100) for the semester (8 total nodes on GCP). Each group is assigned a single static external IP address for SSH access.
Course questionnaire (Week 1): a short onboarding questionnaire released in the first lecture; enrollment/compute allocations will be finalized by the end of the second week.
Students will gain hands-on experience with industry-standard tools and frameworks throughout the course:
Everything in one place: lecture plan, dated timeline, deadlines/milestones, and recommended readings. Dates below are for Fall 2026. Lecture plan subject to change.
| Date | Lecture / Events | Recommended Readings / Resources | Notes |
|---|---|---|---|
| Part 1: Architecture. Defining the Computational Graph and Scaling Strategies | |||
| Aug 26 |
L1. Course Overview and the Modern AI Stack
Key topics
|
||
| Aug 31 |
L2. All About Performance (Part 1)
Key topics
Also:
|
|
|
| Sep 2 |
L3. All About Performance (Part 2)
|
β | |
| Sep 7 |
Labor Day No class
|
β | |
| Sep 9 |
L4. Architectures to Break Bottlenecks: MoE, Sparse & Long-Context Architectures
|
||
| Sep 14 |
L5. Parallelism Strategies
Key topics
|
||
| Sep 16 |
L5. Parallelism Strategies (continued)
L6. FlashAttention + Sonic MoE
|
||
| Sep 21 |
L7. NeMo AutoModel
|
||
| Part 2: Pre-Training of Language Models. The Engineering of βBase Modelsβ | |||
| Sep 23 |
L8. Introduction to Pre-Training
Key topics
Also:
|
β | |
| Sep 28 |
L9. Pre-Training (Part 2): Data for Pre-Training
Key topics
Also:
|
β | |
| Sep 30 |
L10. Modern Architectures + Modern Optimizers
Key topics
|
β | |
| Oct 5 |
L11. MARS Lecture
|
β | |
| Oct 7 |
L12. Case Study: The Pre-Training of Nano-V3
Also:
|
β | |
| Part 3: Post-Training of Language Models. Specializing Models Through SFT and RL | |||
| Oct 12 |
L13. Intro To the LLM Post-Training Lifecycle and Evaluation
Key topics
|
β | |
| Oct 14 |
L14. The Data Powering Post-Training: SFT Data Engineering and RL Environments + Details in Specific Domains
Key topics
|
β | |
| Oct 19 |
L15. Building SFT Datasets: Generation, Tooling, and NeMo Data Designer
|
β | |
| Oct 21 |
L16. Using the NeMo RL Stack for RL Post-Training + NeMo Gym
Also:
|
β | |
| Oct 26 |
L17. Case Study: Post-Training of Nemotron-NanoV3
|
β | |
| Part 4: Efficient Inference. Deployment Preparation and High-Performance Frameworks | |||
| Oct 28 |
L18. Fundamentals and Overview of High-Performance Inference Frameworks
Key topicsA systems overview of modern inference engines, batching/scheduling, memory management, and throughput/latency tradeoffs. |
β | |
| Nov 2 |
L19. Efficient Inference (Part 2): Inference Engines
|
β | |
| Nov 4 |
L20. High-Performance Inference in Practice: Dynamo, TRT-LLM, vLLM, and SGLang
Key topicsServing engines and control runtimes for LLM applications: throughput optimization, KV cache management, and efficient scheduling. Also:
|
β | |
| Nov 9 |
L21. Grok LPX
Also:
|
β | |
| Nov 11 |
Veterans Day No class
|
β | |
| Part 5: LLM Applications and Use Cases. Building Real-World AI Systems | |||
| Nov 16 |
L22. Fundamentals of Context Engineering
Key topicsDesigning prompts, retrieval, memory, and tool usage under latency/cost constraints and reliability targets. |
β | |
| Nov 18 |
L23. Harness Design
Key topicsMethods for engineering software harnesses for modern LLMs. |
β | |
| Nov 23 |
L24. Safety and Security (Part 1): Deployment Guardrails
|
β | |
| Nov 25 |
Non-Instructional Day No class
|
β | |
| Nov 30 |
L25. Safety and Security (Part 2): Open-Weight Safety
|
β | |
| Part 6: Research Greenfields. Emerging Research Directions | |||
| Dec 2 |
L26. Greenfields in Performance & Accelerators
Also:
|
β | |
| Dec 7 |
L27. Greenfields in Pre-Training and Post-Training
Also:
|
β | Peer review cycle |
| Dec 9 |
L28. Greenfields in Inference
Also:
|
β | |
| Dec 10 |
Peer review due Milestone
Submit on Gradescope
|
β | |
| Dec 11 |
Poster session Milestone
No formal lecture
|
β | |
| Dec 16 |
Final deliverables due Milestone
Code + final report
|
β | |
| Dec 18 |
Impact update Milestone
Proof of impact document
|
β | |
This course has no exams. Evaluation is based on assignments, a semester-long research project, an explainer video, and participation.
| Component | Weight |
|---|---|
| Research project | 50% |
| Assignments | 35% |
| Explainer video | 10% |
| Attendance & participation | 5% |
This course requires sustained weekly engagement (group work, compute usage, and in-person checkoffs). It must be taken for the maximum number of units. Letter grade only (no Satisfactory/Non-Satisfactory).
There are four group-based assignments, each spanning roughly 3β4 weeks. Assignments include conceptual questions and hands-on experiments (implementation, profiling, scaling, analysis) with an in-person oral checkoff/presentation per assignment.
All assignment release/due dates are listed in the Consolidated Schedule.
You have 3 total slack days across the semester for assignments. At most one slack day may be applied to any single assignment (extends deadline by 24 hours). No other late work is accepted without an approved exception.
The research project is group-based and open-ended, with an emphasis on producing something useful to the community. Expect 1β2 project check-ins with course staff over the course of the semester.
All project milestones and dates are listed in the Consolidated Schedule.
Each student records a short explainer video on a topic from the course: record yourself at a whiteboard giving a clear, simple explanation of a concept from lecture, where you go further than we had time to.
These are made for your classmates, so make them as helpful as you can. Full credit requires going into more depth on the topic than lecture does, not restating it. Each video is due a month after its topic is introduced in lecture. No video editing is required.
Attendance is required. Participation includes contributing to in-class discussions and project check-ins, posting/answering on Ed, and providing constructive comments on lecture notes and slides.
Collaboration is encouraged within your group. Do not copy code or reports from other groups. If you use external code or AI assistants, cite the source clearly and ensure you understand the work you submit.
This course uses shared GPU infrastructure. For access/quota/networking issues, post on Ed and include timestamps, job IDs, logs, and a short repro description. Multi-node experiments require coordinating with another team to share nodes for a bounded window of time.
Submission platform (Gradescope): all written and code submissions are via Gradescope (entry code posted on Ed), unless explicitly stated otherwise. Oral presentations/checkoffs require sign-up and in-person attendance.