BrainBank

CS 329Z: Engineering AI Agents

9/7/2026, 9:13:56 AM · Source

#ai-agents#tool-use#step-by-step#rag#evaluation#compound-ai#dspy

This course covers the full engineering spectrum of compound AI systems and autonomous agents, from building foundational components like RAG and tool-use from scratch to evaluation, optimization, and frameworks like DSPy.

Welcome!

The shift from monolithic language models to compound AI systems — systems with multiple interacting components including LLMs, retrievers, tools, and optimizers — represents a fundamental change in how AI applications are built for people. This course teaches students how to engineer agentic systems: the full spectrum from simple LLM pipelines to compound AI systems to autonomous agents. Students will learn to pick what types of problems to focus on, decompose problems, select appropriate components, collect and curate data, build evaluations, and reason about the design tradeoffs that arise when building these systems in practice.

Students first build core components (RAG, tool use, agent loops) from scratch, then learn how frameworks like DSPy abstract these patterns. Through two fully applied homework assignments and a quarter-long project, students gain hands-on experience building, optimizing, and evaluating agentic systems.

Class Schedule

Note: the schedule is tentative and subject to change. Classes meet Mondays and Wednesdays, 1:30–2:50 p.m., in Packard 101. Lecture materials will be linked here as they are released.

Week

Date

Lecture

Course Material

1

Wed Sep 23

Foundations & LandscapeIntroduction — What Are Agentic Systems?The spectrum from monolithic models to compound AI systems to agents; when compound systems win; the three engineering challenges (decomposition, data, evaluation); course logistics.

Readings:

Additional readings

2

Mon Sep 28

LLMs for BuildersAPIs & SDKs (litellm), structured I/O and constrained generation, decoding strategies and test-time compute, context engineering, model selection, and cost/latency tradeoffs.

Readings:

Additional readings

2

Wed Sep 30

Building BlocksRetrieval-Augmented Generation (RAG)Grounding and hallucination, embeddings and vector stores, chunking strategies, hybrid search, cross-encoders and late interaction (ColBERT). Hands-on: build a RAG pipeline from scratch.

Readings:

Additional readings

3

Mon Oct 5

Tool Use & Function CallingThe REPL, function-calling APIs, the Model Context Protocol (MCP), designing good tools, code-execution sandboxes, error handling and retries. Hands-on: build a tool-using system from scratch.

Readings:

3

Wed Oct 7

Frameworks & Agent DesignFrameworks & OrchestrationDSPy (signatures, modules, optimizers), LangChain/LangGraph, LlamaIndex; what frameworks abstract vs. what you built from scratch; choosing the right level of abstraction.

Readings:

4

Mon Oct 12

Agent Design Patterns & ScaffoldsThe workflows-vs-agents taxonomy, five composable workflow patterns, agent patterns (ReAct, plan-and-execute, reflection), and scaffolds as design decisions.

Readings:

4

Wed Oct 14

Memory & Multi-Agent SystemsAgent Memory ArchitecturesShort- vs. long-term memory, memory as tool-based actions, the file system as externalized memory, structured memory paradigms, and cross-agent memory.

Readings:

Additional readings

5

Mon Oct 19

Multi-Agent SystemsSingle vs. multi-agent architectures, orchestration patterns, handoffs and state transfer, delegation and collaboration patterns, and the challenges of coordination and error propagation.

Readings:

Additional readings

5

Wed Oct 21

OptimizationOptimizationThe landscape from prompts to fine-tuning; prompt optimization (GEPA, MIPROv2, OPRO, TextGrad); test-time compute scaling; LoRA/QLoRA; distillation; RLHF/DPO at a high level; and when to optimize prompts vs. weights vs. inference compute.

Readings:

Additional readings

6

Mon Oct 26

📺 Guest Lecture (TBA)

6

Wed Oct 28

Data for Agentic SystemsWhat Data Do Agents Need?Traces, demonstrations, and feedback; data for optimization vs. evaluation; data flywheels; synthetic data generation; collecting data from human-agent interaction.

Readings:

Additional readings

7

Mon Nov 2

Data Selection & QualityFinding maximally informative data, filtering and selection strategies, tiny-but-targeted benchmarks, annotation practices, quality assessment, and building datasets from agent traces.

Readings:

Additional readings

7

Wed Nov 4

Evaluation for Agentic SystemsEvaluation Fundamentals & Benchmark DesignWhy evals are hard, the 4-tuple framework (request, environment, stopping criteria, scorer), designing each component, properties of good benchmarks, realistic scaffolding, and reliability dimensions.

Readings:

Additional readings

8

Mon Nov 9

LLM-as-Judge & Evaluation InfrastructureThe three grader types, designing judge prompts, known biases, pairwise vs. pointwise evaluation, non-determinism metrics (pass@k vs. pass^k), harness design, and Anthropic's 8-step roadmap.

Readings:

Additional readings

8

Wed Nov 11

SafetyAgent Safety & GuardrailsPrivacy risks of tool access, prompt injection (including indirect injection), red-teaming, sandboxing and permission models, output guardrails, liability considerations, responsible deployment, and human-in-the-loop patterns.

Readings:

Additional readings

9

Mon Nov 16

📺 Guest Lecture (TBA)

9

Wed Nov 18

Coding Agents & Proactive AgentsCoding & Software AgentsHow coding agents work end-to-end; SWE-agent, Claude Code, and OpenHands architectures; scaffolds as design decisions; SWE-bench and the 4-tuple framework in practice; the future of software development with agents.

Readings:

Additional readings

10

Mon Nov 23

No Class — Thanksgiving Recess

10

Wed Nov 25

No Class — Thanksgiving Recess

11

Mon Nov 30

Proactive AgentsFrom reactive to proactive; General User Models (GUM); Next Action Prediction; open-source proactive personal agents; privacy and trust implications; and when agents should initiate vs. wait (mixed initiative).

Readings:

Additional readings

11

Wed Dec 2

Open Problems & Final DemosFrontiers & Open ProblemsMultimodal agents, web agents and computer use, science agents, long-running agent architectures, production and observability (tracing, monitoring, cost management), and open problems in reliability, scalability, and interpretability.

Additional readings

—

Finals Week (Dec 7–11)

Final Project Demo DayHeld during the end-quarter examination period; exact time and location TBA.

Deadlines

All times are Pacific. Deadlines are subject to change; any changes will be announced in class and posted here.

Week

Deadline

Date

Time

3

HW1 Released

Mon Oct 5

—

3

Project Proposal Due

Fri Oct 9

11:59 p.m.

6

HW2 Released

Mon Oct 26

—

6

HW1 Due

Fri Oct 30

11:59 p.m.

7

Midpoint Demo

Wed Nov 4

In class

7

Midway Report Due

Fri Nov 6

11:59 p.m.

8

Paper Video Due (10 min)

Fri Nov 13

11:59 p.m.

9

HW2 Due

Fri Nov 20

11:59 p.m.

11

Peer Reviews Due (3 videos)

Mon Nov 30

11:59 p.m.

Finals

Final Submission & Final System Demo

Dec 7–11

TBA

Coursework

Two fully applied homework assignments build on the components covered in lecture. Each homework is followed by a 10-minute HW-based quiz where students explain their design decisions and tradeoffs and demonstrate understanding.

  • HW1: Build an Agentic System (Weeks 3–6). Given a repository of research papers, build an agent that answers science questions by retrieving relevant papers and reasoning over them. Part A builds the agent from scratch with litellm (RAG + tool use + an agent loop with a reasoning pattern such as ReAct); Part B rebuilds key components with DSPy and reflects on what the framework abstracts.
  • HW2: Evaluate an Agent (Weeks 6–9). Given a pre-built agent, design a comprehensive evaluation suite with code-based graders, at least one LLM-as-judge eval, benchmark tasks built with the 4-tuple framework (request, environment, stopping criteria, scorer), and error analysis.

Paper Video & Peer Reviews

Each student records a 10-minute video about a recent agent paper of their choosing (due Week 8), then watches and reviews three videos from other students (due after Thanksgiving).

  • Video (7%). 2% selection of a substantive recent paper and a solid explanation of it; 2% your own critique or insight — what you agree/disagree with, limitations, or an interesting question it raises; 2% added value, e.g. reproduce a result, run a small experiment, compare with another method, demo an implementation, or connect it to a real agent; 1% clear and engaging presentation.
  • Peer reviews (3%). 1% each for specific, thoughtful feedback that goes beyond "good presentation" — identify one strength, one weakness or question, and one concrete suggestion.

Grading

  • Project [50%]
    • Proposal [5%]
    • Midway report [5%]
    • Midpoint demo [7%]
    • Final submission [15%]
    • Final system demo [18%]
  • Homework [20%]
    • HW1: Build an Agentic System [10%]
    • HW2: Evaluate an Agent [10%]
  • HW-based quizzes [15%]
    • Quiz 1 (after HW1) [7.5%]
    • Quiz 2 (after HW2) [7.5%]
  • Paper video & peer reviews [10%]
    • Paper video [7%]
    • Peer reviews (3 × 1%) [3%]
  • Participation [5%]
    • In-class discussion, project teamwork, and recitations

Course Project

Students work in groups on a quarter-long project on the theme "Making Life at Stanford Better with Agents." The goal is to build an agentic system that helps with some aspect of Stanford life. Example project ideas:

  • A syllabus reader that extracts deadlines and adds them to your calendar
  • A course-schedule optimizer for next quarter
  • A research-paper discovery and summarization agent
  • A campus-event aggregator and recommender

Milestones

  • Week 3: Project proposal (1-page)
  • Week 6: Midway report (1-page) + midpoint demo — a working prototype is expected
  • Week 10: Final submission (1-page) + final system demo (Demo Day)

Reports are brief (1–2 pages of writing, with an appendix for required structured content such as examples of your agent's failure modes).

Learning map

Stage 1: Foundations & Core Components

  • Foundations of Compound AI: Understand the shift from monolithic language models to multi-component agentic systems and their engineering challenges.
  • LLM APIs & Context Engineering: Master structured outputs, decoding strategies, test-time compute, and handling token context efficiently.
  • Retrieval-Augmented Generation (RAG): Build grounded retrieval mechanisms using vector stores, chunking strategies, and hybrid search.

Stage 2: Orchestration & Advanced Design

  • Tool Use & Function Calling: Implement function-calling APIs, code execution sandboxes, and the Model Context Protocol (MCP).
  • Agentic Frameworks: Explore abstractions like DSPy, LangGraph, and LlamaIndex for compiling declarative workflows.
  • Memory & Multi-Agent Systems: Design short-term and long-term memory architectures along with multi-agent coordination patterns.

Stage 3: Optimization & Evaluation

  • Prompt & Weight Optimization: Scale test-time compute and apply prompt evolution or fine-tuning techniques.
  • Agent Evaluation & LLM-as-Judge: Build rigorous benchmarking environments using the 4-tuple framework and automated scoring graders.
  • Safety & Production Guardrails: Mitigate security threats like prompt injection and implement privacy-preserving guardrails.

Get hands-on — step by step

  1. Set up your Python environment and install the LiteLLM library along with your preferred API keys for model interaction.
  2. Write a script to query an LLM using structured I/O and constrained generation schemas.
  3. Build a basic RAG pipeline by chunking a set of documents, generating embeddings, and implementing a vector similarity search.
  4. Implement a custom tool-calling loop where an LLM can invoke a local REPL or calculator function based on user requests.
  5. Assemble these pieces into a complete ReAct agent loop that reasons about a problem, retrieves data, uses a tool, and loops until completion.

Top 3 sources

  1. 1
    Anthropic: Building Effective Agents

    An engineering guide detailing practical patterns, workflows, and best practices for building robust LLM agents.

    https://www.anthropic.com/engineering/building-effective-agents

  2. 2
    DSPy GitHub Repository

    The official repository for DSPy, a framework for programming—rather than prompting—language models.

    https://github.com/stanfordnlp/dspy

  3. 3
    Model Context Protocol Specification

    The official specification for connecting AI models securely to data sources and development tools.

    https://modelcontextprotocol.io/specification/2025-06-18

Links are AI-suggested — worth a quick sanity check before diving in.