LangGraph Track  /  Module 05  /  Interview Gauntlet
Module 5 of 5 40+ questions
LangGraph · Module 05 · Interview Prep

The Interview Gauntlet
40+ questions, said out loud

No new concepts here, just pressure. This is the separate interview bank: rapid-fire flashcards, deep questions by topic, system-design scenarios, and the gotchas that catch people who only read. Cover the answer, say yours, then check. The starred questions are the ones you must be able to answer cold.

drill in 20-min blocks covers Modules 01-04 recall, not recognition ★ = must-know cold
0

How to drill this so it sticks

Reading answers feels productive and teaches you almost nothing. The only thing that survives interview nerves is active recall. So work this bank in one discipline: read the question, say your full answer aloud as if a person is across the table, then open the card and compare. If your spoken answer missed the key term, that card is not done. Come back to it tomorrow. A few honest passes beat ten lazy re-reads.

The one-line spine

If you can say this sentence and unpack each clause, you can carry most of a LangGraph interview: LangGraph models an agent as a graph of nodes over shared state, where a checkpointer persists that state after every step, which is what unlocks memory, human-in-the-loop, durability, and time travel. Almost every question is a zoom into one clause of that sentence.

1

Rapid-fire flashcards

One term, one crisp definition. Glance at the term, define it aloud, then read. These are the vocabulary you must own.

State
The shared, typed object every node reads and writes. The agent's memory for one run.
Node
A function: takes state, returns a partial update. Where work happens.
Edge
Wiring deciding the next node. Normal = fixed; conditional = router function.
Reducer
Rule merging a node's update into a field. Default overwrite; add appends.
Checkpointer
Backend saving a state snapshot after every step. Short-term memory + durability.
Thread
An independent timeline of state, keyed by thread_id. One conversation.
Store
Long-term memory across threads, keyed by namespace (e.g. user id).
MessagesState
Built-in state with a messages list that appends via a reducer.
Command
A node return carrying update + goto: change state and route in one move.
interrupt()
Durably pause mid-node for a human; resume with Command(resume=...).
Send
Dynamic parallel fan-out (map-reduce): one worker per item, reducer merges results.
Subgraph
A compiled graph used as a node in a bigger graph. Reuse + isolation.
RetryPolicy
Per-node auto-retry on transient failure. Part of durable execution.
create_react_agent
Prebuilt tool-calling agent (think → call tool → observe). Returns a compiled graph.
Supervisor
Central agent that routes to specialist workers one at a time.
Swarm
Peers that hand off directly to each other, no central router.
LangSmith
Tracing/eval tool: see every node, prompt, tool call, token, and latency per run.
Durable execution
Crash mid-run resumes from the last checkpoint instead of restarting.
2

Fundamentals

F1What is LangGraph and what problem does it solve?
A low-level orchestration framework for building stateful, long-running agents as graphs: nodes (steps) over shared state, connected by edges. It solves the problem that real agents need branching, loops, memory, pausing, and recovery, which become an unmaintainable tangle of conditionals in plain code. Modelling the agent as a graph makes control flow explicit and lets the runtime persist, resume, stream, and pause.
F2LangGraph vs LangChain, vs a plain while-loop?
LangChain provides building blocks (models, tools, prompts); LangGraph provides orchestration, deciding what runs next and carrying state, and can be used without LangChain. Versus a plain while-loop, the graph makes branches and cycles explicit and, because each step returns an update instead of mutating state, the runtime can checkpoint every step, which a raw loop cannot do, giving you persistence, HITL, and recovery for free.
F3Walk through the execution model from START to END.
The runtime begins at START, follows the edge to the first node, runs it, merges its returned update into state (via the field's reducer), checkpoints, then re-reads the edges to find the next node, repeating until an edge leads to END. Conditional edges run a router over current state to pick the next node; cycles are allowed.
F4What exactly does a node receive and return, and why not mutate state?
It receives the current state and returns a dict of only the fields it changes. It must not mutate state in place because the runtime needs to apply updates through reducers and snapshot each step; returning an immutable update is what makes checkpointing, parallel merges, and time travel possible.
F5Normal edge vs conditional edge, with the API.
Normal: add_edge(A, B), an unconditional jump. Conditional: add_conditional_edges(A, router), where router(state) returns the name of the next node (or END). Conditional edges express branching and loops.
F6What does .compile() do?
Turns the builder description into a runnable graph: validates wiring and returns an object with invoke/stream. It is also where you attach cross-cutting features like a checkpointer and store.
3

State & memory

S1What is a reducer and why does parallelism require one?
A reducer merges a node's update into a field. Default is overwrite; Annotated[list, add] appends. Parallel nodes writing the same field would conflict, the runtime cannot guess how to combine two writes, so a reducer is required to merge them deterministically. It is also how accumulation (findings, messages) works.
S2How does an agent remember a conversation across calls?
Compile with a checkpointer (saves state after each step) and pass a thread_id in config on each call. Same thread id loads the last checkpoint and continues; a new thread id starts blank. In production the checkpointer is backed by SQLite or Postgres.
S3Short-term vs long-term memory.
Short-term = checkpointer, scoped to a thread (this conversation). Long-term = a Store, keyed by a namespace like user id, holding facts that persist across threads (preferences, profile). "This chat" vs "this user, always".
S4Why does attaching a checkpointer force you to pass a thread_id?
The checkpointer is shared storage for many conversations. The thread id selects which timeline to read and write; without it the runtime cannot locate the right state and errors.
S5Streaming modes and when to use each.
"updates" = each node's change (progress UI); "values" = full state per step (debugging); "messages" = LLM tokens as they stream (typewriter chat UI). Stream when a run is long enough that a blank wait feels broken.
S6What is time travel and what enables it?
Because the checkpointer snapshots after every step, a thread is a full history. get_state_history lists snapshots; you can inspect or resume from any past checkpoint (optionally edited). Same machinery underlies HITL and conversation forking.
S7When would you use separate input/output schemas?
When the internal state carries scratch fields you do not want callers to see or send. You pass input_schema and output_schema to keep the public contract clean while the graph keeps private working fields.
4

Control flow & human-in-the-loop

C1Explain human-in-the-loop end to end.
A node calls interrupt(payload); the run durably pauses (state is already checkpointed) and the payload surfaces to your app. You resume with Command(resume=value) on the same thread; the node re-runs up to the interrupt and continues with the human's value. Needs a checkpointer and thread id. Flavours: approve/reject, edit state, review a tool call, ask for input, all the same mechanism.
C2Command vs conditional edge.
Command lets a node return an update and a goto together, routing itself; a conditional edge keeps routing in a separate function. Use Command when the decision and the work are intertwined, or to route plus hand off to another agent in one step.
C3Why must side effects come after interrupt()?
On resume the node re-executes from the top to the interrupt, so code before it runs twice. Put sends, charges, and writes after the interrupt so they fire exactly once.
C4What is Send and how does map-reduce work here?
A node returns a list of Send(node, input), one per item (the map); the runtime runs that node in parallel for each, and a reducer on the result field merges all outputs (the reduce). The number of workers is decided at runtime.
C5What is a subgraph and how do parent and child share state?
A compiled graph used as a node in a larger graph. They share data through common state keys; if schemas differ, wrap the subgraph in a function node that translates state in and out.
C6What makes execution durable, and why does it matter?
Every step is checkpointed (resume from last good step on crash) and nodes can carry a RetryPolicy (auto-retry transient failures). It matters because real agents call flaky APIs and run long; durability stops one hiccup from losing the whole task.
5

Multi-agent & production

M1When do you reach for multiple agents instead of one?
When one agent's toolset or responsibilities grow enough that prompt-following degrades, or sub-tasks need different instructions, models, or independent testing. Otherwise prefer a single agent, multi-agent buys modularity at the cost of coordination, latency, and tokens.
M2Supervisor vs swarm vs agent-as-tool.
All answer "who decides routing." Supervisor: a central agent routes to workers that report back. Swarm: peers hand off directly (a Command with graph=Command.PARENT). Agent-as-tool: a sub-agent is wrapped as a tool the caller invokes. All use the same Command routing and shared state.
M3What is create_react_agent under the hood?
A prebuilt compiled graph running the ReAct loop (think → call tool → observe → think). Because it is a normal graph it supports checkpointers, streaming, and interrupts, and drops into a larger graph as a node or subgraph.
M4How do you debug a flaky agent in production?
Read the LangSmith trace for a failing run: each node, the exact prompt, every tool input/output, plus token and latency per step, to pinpoint the bad node or tool call. Combine with checkpoint history (time travel) to inspect state at each step.
M5What changes when you go to production?
In-memory checkpointer becomes Postgres-backed (durable, shared memory); the graph runs as a service (LangGraph Platform, self-hosted server, or embedded in your API) handling concurrency/scaling; add LangSmith tracing and Studio for visual debugging.
M6How does a tool get described to the model?
A tool is a function; its name, typed signature, and docstring become the description the model reads to decide when and how to call it. Clear docstrings and argument names directly improve tool-selection accuracy.
6

System-design scenarios

Open-ended. Talk through the design out loud; the card is a strong reference answer, not the only one.

D1Design a customer-support agent that can issue refunds, with a human approving any refund over $100.
State holds the conversation (MessagesState) plus a proposed action. A ReAct agent handles chat and calls tools (lookup order, draft refund). Route the proposed refund through a node: if amount > $100, call interrupt() to pause for human approval, otherwise proceed. Keep the actual refund side effect after the interrupt. Compile with a Postgres checkpointer so the pause survives and each customer is a thread. Add LangSmith tracing. Mention the re-run gotcha as why the side effect goes after the interrupt.
D2Design a research agent that investigates N sub-topics in parallel and writes one report.
A planner node produces the list of sub-topics, then returns Send("research_one", {topic}) for each, fanning out in parallel. The findings field uses an add reducer so all parallel results merge into one list. A synthesizer node reads merged findings and writes the report. Add a checkpointer for resumability and stream "updates" so the user sees progress. This is textbook map-reduce with Send + reducer.
D3Design a coding assistant: a supervisor delegating to a planner, a coder, and a reviewer.
Supervisor pattern. Each specialist is a node or subgraph (the coder may itself be a ReAct agent with file tools). The supervisor routes with Command(goto=...) based on state (plan ready? code written? review passed?); workers return Command(update=..., goto="supervisor"). Loop coder↔reviewer until the reviewer approves, then END. Add a checkpointer and optionally an interrupt() before applying changes. Justify multi-agent: distinct instructions/tools per role and independent testing.
D4A long-running data-pipeline agent must survive server restarts and not redo finished work.
Lean on durable execution. Break the pipeline into nodes so each completed step is checkpointed; back the checkpointer with Postgres. On restart, re-invoke on the same thread_id and the runtime resumes from the last good checkpoint. Add RetryPolicy to flaky external-call nodes. Make node side effects idempotent where possible, since a node can re-run. This is the "why LangGraph for long-running work" answer: persistence.
7

Gotchas & traps

The questions that separate "read a tutorial" from "actually built one." Each is a real mistake people make.

T1"My node returns the full state and things break / get slow." Why?
A node should return only the fields it changed, not a full copy. Returning everything can clobber reducer-managed fields (e.g. re-sending the whole messages list makes the append reducer duplicate). Return the minimal update.
T2"interrupt() throws or never pauses." What did I forget?
A checkpointer. interrupt() needs persisted state to pause and resume, so the graph must be compiled with a checkpointer and invoked with a thread_id. No checkpointer, no durable pause.
T3"My email got sent twice after a human approval." Why?
The side effect was placed before interrupt() in the node. On resume the node re-runs from the top to the interrupt, so the send executed twice. Move all side effects after the interrupt.
T4"Two parallel nodes write the same field and the run errors." Fix?
That field has no reducer, so the runtime cannot merge concurrent writes. Annotate it with a reducer (e.g. Annotated[list, add]) so the parallel updates combine deterministically.
T5"My agent forgets everything between turns even with MessagesState." Why?
MessagesState only appends within a single run. Cross-call memory needs a checkpointer plus a consistent thread_id. Without them each invoke starts from the input you pass, with no saved history.
T6"I jumped straight to a 5-agent system and it's a mess." What's the lesson?
Multi-agent is not free, it adds coordination, latency, tokens, and failure points. Start with one well-prompted agent and split only when toolset size or distinct sub-tasks justify it. Interviewers respect "I'd start single and measure before adding agents."
T7"My loop never terminates." What's the usual cause?
The conditional edge / router never returns END; the exit condition is never met (a counter not incremented, a flag never set). Always have a clear termination branch, and consider a recursion/step limit as a safety net.
8

One last gut check

Before you call yourself interview-ready, try this without notes: in two minutes, explain to an imaginary interviewer how you would build a research assistant that holds a conversation, pauses for human approval before sending anything, researches several topics in parallel, and runs reliably in production. If you can narrate Sahil's whole agent, state and reducers, checkpointer and threads, interrupt and Send, supervisor and tracing, in one flowing answer, you have actually learned this, not just read it.

You are done when

You can answer every ★ question cold, narrate the two-minute design above, and explain at least three of the seven traps from having understood why, not from memorising the card. That is the bar. Go get the offer.