CS 329Z: Engineering AI Agents
9/7/2026, 9:13:56 AM · Source
This course covers the full engineering spectrum of compound AI systems and autonomous agents, from building foundational components like RAG and tool-use from scratch to evaluation, optimization, and frameworks like DSPy.
Welcome!
The shift from monolithic language models to compound AI systems — systems with multiple interacting components including LLMs, retrievers, tools, and optimizers — represents a fundamental change in how AI applications are built for people. This course teaches students how to engineer agentic systems: the full spectrum from simple LLM pipelines to compound AI systems to autonomous agents. Students will learn to pick what types of problems to focus on, decompose problems, select appropriate components, collect and curate data, build evaluations, and reason about the design tradeoffs that arise when building these systems in practice.
Students first build core components (RAG, tool use, agent loops) from scratch, then learn how frameworks like DSPy abstract these patterns. Through two fully applied homework assignments and a quarter-long project, students gain hands-on experience building, optimizing, and evaluating agentic systems.
Class Schedule
Note: the schedule is tentative and subject to change. Classes meet Mondays and Wednesdays, 1:30–2:50 p.m., in Packard 101. Lecture materials will be linked here as they are released.
Week
Date
Lecture
Course Material
1
Wed Sep 23
Foundations & LandscapeIntroduction — What Are Agentic Systems?The spectrum from monolithic models to compound AI systems to agents; when compound systems win; the three engineering challenges (decomposition, data, evaluation); course logistics.
Readings:
Additional readings
2
Mon Sep 28
LLMs for BuildersAPIs & SDKs (litellm), structured I/O and constrained generation, decoding strategies and test-time compute, context engineering, model selection, and cost/latency tradeoffs.
Readings:
Additional readings
2
Wed Sep 30
Building BlocksRetrieval-Augmented Generation (RAG)Grounding and hallucination, embeddings and vector stores, chunking strategies, hybrid search, cross-encoders and late interaction (ColBERT). Hands-on: build a RAG pipeline from scratch.
Readings:
Additional readings
3
Mon Oct 5
Tool Use & Function CallingThe REPL, function-calling APIs, the Model Context Protocol (MCP), designing good tools, code-execution sandboxes, error handling and retries. Hands-on: build a tool-using system from scratch.
Readings:
3
Wed Oct 7
Frameworks & Agent DesignFrameworks & OrchestrationDSPy (signatures, modules, optimizers), LangChain/LangGraph, LlamaIndex; what frameworks abstract vs. what you built from scratch; choosing the right level of abstraction.
Readings:
4
Mon Oct 12
Agent Design Patterns & ScaffoldsThe workflows-vs-agents taxonomy, five composable workflow patterns, agent patterns (ReAct, plan-and-execute, reflection), and scaffolds as design decisions.
Readings:
4
Wed Oct 14
Memory & Multi-Agent SystemsAgent Memory ArchitecturesShort- vs. long-term memory, memory as tool-based actions, the file system as externalized memory, structured memory paradigms, and cross-agent memory.
Readings:
Additional readings
5
Mon Oct 19
Multi-Agent SystemsSingle vs. multi-agent architectures, orchestration patterns, handoffs and state transfer, delegation and collaboration patterns, and the challenges of coordination and error propagation.
Readings:
Additional readings
5
Wed Oct 21
OptimizationOptimizationThe landscape from prompts to fine-tuning; prompt optimization (GEPA, MIPROv2, OPRO, TextGrad); test-time compute scaling; LoRA/QLoRA; distillation; RLHF/DPO at a high level; and when to optimize prompts vs. weights vs. inference compute.
Readings:
Additional readings
6
Mon Oct 26
📺 Guest Lecture (TBA)
6
Wed Oct 28
Data for Agentic SystemsWhat Data Do Agents Need?Traces, demonstrations, and feedback; data for optimization vs. evaluation; data flywheels; synthetic data generation; collecting data from human-agent interaction.
Readings:
Additional readings
7
Mon Nov 2
Data Selection & QualityFinding maximally informative data, filtering and selection strategies, tiny-but-targeted benchmarks, annotation practices, quality assessment, and building datasets from agent traces.
Readings:
Additional readings
7
Wed Nov 4
Evaluation for Agentic SystemsEvaluation Fundamentals & Benchmark DesignWhy evals are hard, the 4-tuple framework (request, environment, stopping criteria, scorer), designing each component, properties of good benchmarks, realistic scaffolding, and reliability dimensions.
Readings:
Additional readings
8
Mon Nov 9
LLM-as-Judge & Evaluation InfrastructureThe three grader types, designing judge prompts, known biases, pairwise vs. pointwise evaluation, non-determinism metrics (pass@k vs. pass^k), harness design, and Anthropic's 8-step roadmap.
Readings:
Additional readings
8
Wed Nov 11
SafetyAgent Safety & GuardrailsPrivacy risks of tool access, prompt injection (including indirect injection), red-teaming, sandboxing and permission models, output guardrails, liability considerations, responsible deployment, and human-in-the-loop patterns.
Readings:
Additional readings
9
Mon Nov 16
📺 Guest Lecture (TBA)
9
Wed Nov 18
Coding Agents & Proactive AgentsCoding & Software AgentsHow coding agents work end-to-end; SWE-agent, Claude Code, and OpenHands architectures; scaffolds as design decisions; SWE-bench and the 4-tuple framework in practice; the future of software development with agents.
Readings:
Additional readings
10
Mon Nov 23
No Class — Thanksgiving Recess
10
Wed Nov 25
No Class — Thanksgiving Recess
11
Mon Nov 30
Proactive AgentsFrom reactive to proactive; General User Models (GUM); Next Action Prediction; open-source proactive personal agents; privacy and trust implications; and when agents should initiate vs. wait (mixed initiative).
Readings:
Additional readings
11
Wed Dec 2
Open Problems & Final DemosFrontiers & Open ProblemsMultimodal agents, web agents and computer use, science agents, long-running agent architectures, production and observability (tracing, monitoring, cost management), and open problems in reliability, scalability, and interpretability.
Additional readings
—
Finals Week (Dec 7–11)
Final Project Demo DayHeld during the end-quarter examination period; exact time and location TBA.
Deadlines
All times are Pacific. Deadlines are subject to change; any changes will be announced in class and posted here.
Week
Deadline
Date
Time
3
HW1 Released
Mon Oct 5
—
3
Project Proposal Due
Fri Oct 9
11:59 p.m.
6
HW2 Released
Mon Oct 26
—
6
HW1 Due
Fri Oct 30
11:59 p.m.
7
Midpoint Demo
Wed Nov 4
In class
7
Midway Report Due
Fri Nov 6
11:59 p.m.
8
Paper Video Due (10 min)
Fri Nov 13
11:59 p.m.
9
HW2 Due
Fri Nov 20
11:59 p.m.
11
Peer Reviews Due (3 videos)
Mon Nov 30
11:59 p.m.
Finals
Final Submission & Final System Demo
Dec 7–11
TBA
Coursework
Two fully applied homework assignments build on the components covered in lecture. Each homework is followed by a 10-minute HW-based quiz where students explain their design decisions and tradeoffs and demonstrate understanding.
- HW1: Build an Agentic System (Weeks 3–6). Given a repository of research papers, build an agent that answers science questions by retrieving relevant papers and reasoning over them. Part A builds the agent from scratch with litellm (RAG + tool use + an agent loop with a reasoning pattern such as ReAct); Part B rebuilds key components with DSPy and reflects on what the framework abstracts.
- HW2: Evaluate an Agent (Weeks 6–9). Given a pre-built agent, design a comprehensive evaluation suite with code-based graders, at least one LLM-as-judge eval, benchmark tasks built with the 4-tuple framework (request, environment, stopping criteria, scorer), and error analysis.
Paper Video & Peer Reviews
Each student records a 10-minute video about a recent agent paper of their choosing (due Week 8), then watches and reviews three videos from other students (due after Thanksgiving).
- Video (7%). 2% selection of a substantive recent paper and a solid explanation of it; 2% your own critique or insight — what you agree/disagree with, limitations, or an interesting question it raises; 2% added value, e.g. reproduce a result, run a small experiment, compare with another method, demo an implementation, or connect it to a real agent; 1% clear and engaging presentation.
- Peer reviews (3%). 1% each for specific, thoughtful feedback that goes beyond "good presentation" — identify one strength, one weakness or question, and one concrete suggestion.
Grading
- Project [50%]
- Proposal [5%]
- Midway report [5%]
- Midpoint demo [7%]
- Final submission [15%]
- Final system demo [18%]
- Homework [20%]
- HW1: Build an Agentic System [10%]
- HW2: Evaluate an Agent [10%]
- HW-based quizzes [15%]
- Quiz 1 (after HW1) [7.5%]
- Quiz 2 (after HW2) [7.5%]
- Paper video & peer reviews [10%]
- Paper video [7%]
- Peer reviews (3 × 1%) [3%]
- Participation [5%]
- In-class discussion, project teamwork, and recitations
Course Project
Students work in groups on a quarter-long project on the theme "Making Life at Stanford Better with Agents." The goal is to build an agentic system that helps with some aspect of Stanford life. Example project ideas:
- A syllabus reader that extracts deadlines and adds them to your calendar
- A course-schedule optimizer for next quarter
- A research-paper discovery and summarization agent
- A campus-event aggregator and recommender
Milestones
- Week 3: Project proposal (1-page)
- Week 6: Midway report (1-page) + midpoint demo — a working prototype is expected
- Week 10: Final submission (1-page) + final system demo (Demo Day)
Reports are brief (1–2 pages of writing, with an appendix for required structured content such as examples of your agent's failure modes).
Learning map
Stage 1: Foundations & Core Components
- Foundations of Compound AI: Understand the shift from monolithic language models to multi-component agentic systems and their engineering challenges.
- LLM APIs & Context Engineering: Master structured outputs, decoding strategies, test-time compute, and handling token context efficiently.
- Retrieval-Augmented Generation (RAG): Build grounded retrieval mechanisms using vector stores, chunking strategies, and hybrid search.
Stage 2: Orchestration & Advanced Design
- Tool Use & Function Calling: Implement function-calling APIs, code execution sandboxes, and the Model Context Protocol (MCP).
- Agentic Frameworks: Explore abstractions like DSPy, LangGraph, and LlamaIndex for compiling declarative workflows.
- Memory & Multi-Agent Systems: Design short-term and long-term memory architectures along with multi-agent coordination patterns.
Stage 3: Optimization & Evaluation
- Prompt & Weight Optimization: Scale test-time compute and apply prompt evolution or fine-tuning techniques.
- Agent Evaluation & LLM-as-Judge: Build rigorous benchmarking environments using the 4-tuple framework and automated scoring graders.
- Safety & Production Guardrails: Mitigate security threats like prompt injection and implement privacy-preserving guardrails.
Get hands-on — step by step
- Set up your Python environment and install the LiteLLM library along with your preferred API keys for model interaction.
- Write a script to query an LLM using structured I/O and constrained generation schemas.
- Build a basic RAG pipeline by chunking a set of documents, generating embeddings, and implementing a vector similarity search.
- Implement a custom tool-calling loop where an LLM can invoke a local REPL or calculator function based on user requests.
- Assemble these pieces into a complete ReAct agent loop that reasons about a problem, retrieves data, uses a tool, and loops until completion.
Top 3 sources
- 1Anthropic: Building Effective Agents
An engineering guide detailing practical patterns, workflows, and best practices for building robust LLM agents.
https://www.anthropic.com/engineering/building-effective-agents
- 2DSPy GitHub Repository
The official repository for DSPy, a framework for programming—rather than prompting—language models.
https://github.com/stanfordnlp/dspy
- 3Model Context Protocol Specification
The official specification for connecting AI models securely to data sources and development tools.
https://modelcontextprotocol.io/specification/2025-06-18
Links are AI-suggested — worth a quick sanity check before diving in.