Multi-Agent Systems & Production
from one agent to a team, shipped
Sahil's solo agent is doing too much. Here you give it tools, then split it into a team: a supervisor that delegates to specialists. Then the part most tutorials skip, taking it to production: seeing what it does (observability) and getting it running for real users (deployment).
One agent is doing too much
Sahil's single agent now has to plan, search, judge sources, and write, all in one prompt with one growing pile of tools. The prompt is bloated, it picks the wrong tool, and when something goes wrong he cannot tell which job failed. This is the classic moment teams reach for multiple agents: split the one overloaded brain into focused specialists, each with a small toolset and a clear job, coordinated by a graph. The reassuring part: a multi-agent system is still just a LangGraph graph. Each agent is a node (often a subgraph), and the coordination is edges. You already know the primitives.
One agent with ten tools and a paragraph of instructions is unreliable and unobservable. Sahil wants a researcher that only searches, a writer that only writes, and a supervisor that decides who works next, then he wants to ship it and watch it run.
Tools and the prebuilt ReAct agent
Before teams, the single most useful building block: an agent that can call tools. A tool is just a function
the model is allowed to invoke, a web search, a calculator, a database lookup. The loop is "model thinks, maybe calls a tool,
reads the result, thinks again", which is the ReAct pattern you built by hand in Module 01. LangGraph ships it prebuilt as
create_react_agent, so you do not rewire that loop every time.
from langgraph.prebuilt import create_react_agent
def web_search(query: str) -> str:
"""Search the web for a query and return the top result.""" # docstring = tool description the model reads
...
agent = create_react_agent(model=llm, tools=[web_search])
agent.invoke({"messages": [user("Who won the 2022 World Cup?")]})
# Under the hood: the same think → call tool → observe → think loop as Module 01,
# but compiled, batteries-included, and ready to drop into a bigger graph as a node.
create_react_agent returns a compiled graph. It takes a checkpointer, streams, and interrupts just
like everything you built by hand. The prebuilt is a convenience, not a different world. Knowing the manual version (Module 01)
is what lets you debug the prebuilt when it misbehaves.
Why split into multiple agents
The honest answer to "should I use multiple agents?" is often "not yet". One well-prompted agent with a few tools beats a tangle of agents you cannot debug. But there is a real threshold where splitting pays off, and naming it keeps you out of trouble.
Split when
The toolset is too big for one prompt to choose well; jobs need different instructions or models; or you want each part testable and observable on its own.
Stay single when
A few tools and one coherent job. Multiple agents add coordination overhead, more tokens, more latency, and more places to fail. Do not pay that for nothing.
"When would you use a multi-agent architecture?" When a single agent's responsibilities or toolset grow large enough that prompt-following degrades, or when sub-tasks need genuinely different instructions, models, or independent testing. Otherwise prefer one agent: multi-agent buys modularity at the cost of coordination, latency, and tokens.
The supervisor pattern
The most common and most interview-relevant multi-agent shape. A supervisor agent sits at the centre. It
looks at the state, decides which specialist should act next, routes to it, gets the result back, and decides again, until the
job is done. The workers do not talk to each other; they all report to the supervisor. It is an org chart, and it maps cleanly
onto a graph where the supervisor routes with Command (Module 03) and each worker is a node or subgraph.
The supervisor is the only router. Each worker does its job and hands control back. Add a worker = add a node + a route.
from langgraph.types import Command
from typing import Literal
def supervisor(state) -> Command[Literal["researcher", "writer", "__end__"]]:
nxt = pick_next(state) # LLM decides: who works next, or done?
return Command(goto=nxt) # route, same Command from Module 03
def researcher(state):
return Command(update={"findings": [...]}, goto="supervisor") # report back
# Workers always goto "supervisor". The supervisor is the hub of the wheel.
# (langgraph also ships langgraph-supervisor to scaffold this for you.)
You can hand-build the supervisor (best for understanding) or use the langgraph-supervisor library,
which scaffolds the hub-and-workers wiring from a list of agents. Build it by hand once, then reach for the prebuilt.
Handoffs and the swarm pattern
The supervisor is hub-and-spoke. The other shape is the swarm: agents hand off directly to each
other, peer to peer, no central boss. The researcher decides on its own to pass to the writer. The tool that makes this work is
a handoff: a Command whose goto targets another agent, with
graph=Command.PARENT so it can jump to a sibling in the parent graph rather than inside itself.
def handoff_to_writer(state):
return Command(
goto="writer",
graph=Command.PARENT, # jump to a SIBLING in the parent graph
update={"handoff_note": "research done, your turn"},
)
# Supervisor = central router. Swarm = peers handing off directly.
# Both are just Command routing; the difference is who decides.
| Pattern | Who decides routing | Good when |
|---|---|---|
| Supervisor | One central agent | You want oversight, clear control, easy debugging. The default choice. |
| Swarm | Each agent, peer to peer | Specialists with clear "next owner" handoffs; less central bottleneck. |
| Agent-as-tool | The calling agent | A sub-agent is wrapped as a tool another agent can call like any function. |
Supervisor, swarm, and agent-as-tool are not three different technologies. They are three answers to one question:
who decides what runs next? A central node, the peers themselves, or the caller. All three are expressed with the same
Command routing and shared state you already learned. That framing is gold in an interview.
Observability: see what the agent actually did
An agent that works on your laptop and fails silently in production is worthless. You need to see every step: which node ran, what the model was prompted with, what each tool returned, how many tokens it cost, where it got slow. LangChain's tracing tool, LangSmith, captures all of that automatically. You usually turn it on with a couple of environment variables, no code change, and every run becomes a clickable trace of nodes, prompts, tool calls, latencies, and costs.
# set these and your existing graph is traced, no code change
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=... # from smith.langchain.com
# now every invoke/stream appears as a step-by-step trace you can inspect
Trace
Every node, prompt, and tool call for a single run, in order, with inputs and outputs.
Cost & latency
Tokens and time per step, so you can find the slow or expensive node and fix it.
Evals
Run datasets through the agent and score quality over time, so changes are measured, not guessed.
Most "the agent is broken" problems are really "I cannot see what the agent did" problems. Tracing is the difference between debugging by guesswork and debugging by reading the actual run. Interviewers ask how you would debug a flaky agent; "I read the LangSmith trace to find which node or tool call went wrong" is the answer they want.
Deployment: from script to service
Your compiled graph runs fine in a script. Production means it runs as a service: an API many users hit concurrently, with persistence backed by a real database, background runs, and scaling. You have three broad options, and knowing the trade-off is enough for now.
| Option | What it is | Trade-off |
|---|---|---|
| LangGraph Platform | Managed hosting for LangGraph apps: API, persistence, scaling, cron, all provided. | Fastest to production; you run less infrastructure yourself. |
| Self-hosted server | Run the LangGraph server in your own infra (containers) with your own Postgres. | Full control; you own ops, scaling, and upgrades. |
| Embed the graph | Import the compiled graph into your own FastAPI/Flask app and call it. | Maximum flexibility; you rebuild the serving features yourself. |
Whichever you pick, two things carry over from this whole track. First, the checkpointer becomes a real database (Postgres), so memory and durability survive restarts and scale across instances. Second, LangGraph Studio lets you visualise and step through your graph while you build, a visual debugger that draws the exact nodes and edges you wrote.
Notice the whole track converging: state and reducers (M2) define what persists, the checkpointer (M2) becomes production Postgres here, interrupts and durability (M3) keep long runs safe at scale, and multi-agent graphs (M4) are deployed as one service. You did not learn four disconnected topics; you built one thing, layer by layer.
Hands-on: a two-agent team
Build the smallest real team: a supervisor routing between a researcher (one search tool) and a writer. Trace it in LangSmith and read the run. This is the capstone of the build-along track.
from langgraph.graph import StateGraph, START, MessagesState
from langgraph.types import Command
from typing import Literal
def supervisor(state) -> Command[Literal["researcher", "writer", "__end__"]]:
# TODO 1: decide next worker from state; return Command(goto=...)
...
def researcher(state):
# TODO 2: do a (fake) search, return Command(update={...}, goto="supervisor")
...
def writer(state):
# TODO 3: write final answer, return Command(update={...}, goto="supervisor")
...
b = StateGraph(MessagesState)
for n in (supervisor, researcher, writer): b.add_node(n.__name__, n)
b.add_edge(START, "supervisor")
graph = b.compile()
# Run it, then set LANGSMITH_TRACING=true and read the trace.
- The supervisor routed to a worker, got control back, and eventually ended.
- Each worker returned a
Commandthat updated state and went back to the supervisor. - You built a
create_react_agentwith one real tool and watched it call the tool. - You enabled LangSmith and read a full trace of nodes, prompts, and tool calls.
- You can explain supervisor vs swarm as "who decides routing" in one sentence.
Interview check
The full bank is Module 05. = must-know cold.
Q1When should you use multiple agents instead of one?
Q2Explain the supervisor pattern.
Command(goto=...) and workers are nodes or subgraphs.Q3Supervisor vs swarm vs agent-as-tool?
Command with graph=Command.PARENT). Agent-as-tool: a sub-agent is wrapped as a
tool the caller invokes. Same Command routing underneath.