Skip to main content
AI Interview Question
INTERVIEW GUIDEAgents8 questions8 min readOct 4, 2026

LangGraph Agent System Design Interviews: State, Graphs, Checkpoints, and HITL

Interviewers probe durable agent workflows—shared state with reducers, routing and cycles, checkpoints, and human gates. How to design with LangGraph vs a simple ReAct loop.

LangGraph Agent System Design Interviews: State, Graphs, Checkpoints, and HITL

Interviewers for GenAI, applied LLM, and agent-platform roles rarely ask you to recite what LangGraph is. They ask whether you can design a durable, inspectable agent workflow: shared state with clear merge rules, explicit routing (including cycles), checkpointed resume after failure or pause, and safe human gates. The strong answer frames LangGraph as a low-level orchestration runtime for long-running stateful agents—not a chatbot wrapper—and knows when a plain ReAct-style loop is enough.

This guide focuses on architecture trade-offs for agent system design with LangGraph alone. It does not cover MCP or tool-protocol wiring; treat that as a separate tooling layer.

## What interviewers are actually probing

Toy demos hide the hard parts. Production agent platforms fail on resume after crash, on parallel updates stomping state, on human approval that re-runs side effects, and on cycles that never terminate. Graph design interviews test whether you can make those failure modes visible and controllable: explicit state schemas, reducers under fan-out, checkpoints at super-step boundaries, and interrupt patterns that stay idempotent.

## Core model: state, nodes, edges

LangGraph’s durable mental model is small:

- **State** is a shared snapshot for the graph run (often a TypedDict or Pydantic schema). - **Nodes** are units of work. Each node returns a **partial update**, not a full rewrite of the world. - **Edges** define control flow: fixed next hops, conditional routing, or dynamic commands. - Execution is Pregel-style: the runtime advances in **super-steps**. You **compile** the graph before you run it.

If you only remember one sentence for the whiteboard: *nodes emit deltas; reducers (or defaults) merge them into the next shared state.*

## State design and reducers

Interviewers care less about “use TypedDict” and more about **how updates merge**.

By default, a key’s new value **overwrites** the old one. That is fine for scalars and for single-writer paths. Under parallel fan-out—or whenever multiple nodes touch the same list or message history—you need **custom reducers**: for example `Annotated[..., operator.add]` or the messages helper `add_messages`. Reducers define how concurrent or sequential partial updates combine.

A classic pitfall: with a merging list reducer, returning `[]` does **not** clear the list. Empty appends nothing. To replace the accumulated value you need an explicit overwrite mechanism (for example LangGraph’s `Overwrite` pattern), not “set it to empty and hope.” Strong candidates say this out loud before the interviewer traps them.

Design state keys for the job: inputs the agent reads, intermediate drafts, flags for approval, tool results, and a clear signal for termination. Keep the schema typed so resume and time-travel edits stay inspectable.

## Routing and control flow

Three routing styles show up on whiteboards:

1. **Fixed edges** — always go from A to B. 2. **Conditional edges** — a routing function picks the next node from state. 3. **`Command(update=..., goto=...)`** — a node returns both a state update and a dynamic next hop.

Do **not** mix static `add_edge` routing and `Command(goto=...)` from the **same** node. Pick one authority for “where we go next” from that node, or you get ambiguous graphs that are hard to reason about in interviews and in production.

For map-reduce style fan-out, **`Send`** is the usual pattern: spawn parallel work units that write back through reducers. For cycles (retry, re-prompt, tools-until-done), always show an **explicit path to END**. Pair that with `recursion_limit` (and awareness of `GraphRecursionError`). Proactive checks against remaining steps (`RemainingSteps`-style guards) beat hoping the loop exits.

## Persistence and memory

Distinguish short-term and long-term memory clearly:

- **Checkpointers** provide **thread-scoped short-term memory**. They persist checkpoints at **super-step boundaries**. The stable **`thread_id`** is the resume cursor. - **Stores** hold **cross-thread long-term memory** (facts, preferences, shared artifacts across conversations).

For tests, `InMemorySaver` is fine. For production pause/resume and human-in-the-loop, use a **durable checkpointer** (for example Postgres or SQLite). Saying “we’ll use InMemorySaver in prod HITL” is a common fail.

## Human-in-the-loop

HITL is where agent interviews get serious.

- **Dynamic** `interrupt()` vs **static** `interrupt_before` / `interrupt_after` breakpoints: know both. Dynamic interrupts carry a payload; keep that payload **JSON-serializable**. - Resume with `Command(resume=...)` on the **same** `thread_id`, with a checkpointer attached. - On resume, the **node restarts from the top**. Any side effects before the interrupt must be **idempotent** (or moved after approval). Charging a card, sending email, or mutating a ticket before the interrupt without idempotency keys is a red flag. - Avoid `while True` plus multiple `interrupt()` calls inside one node. Prefer a **conditional-edge re-prompt loop** so each interrupt boundary is a clean super-step. - Common patterns: **approval** (refund execute), **review-edit** (human edits draft then continue), **tool-gate** (block risky tools until a human allows them).

Wrapping `interrupt` in a bare `try/except` that swallows it, or reordering/conditionally skipping interrupts inside one node, breaks resume semantics. Call that out before the interviewer does.

## When LangGraph vs a simpler loop

Use a decision table on the whiteboard:

| Situation | Prefer | |-----------|--------| | One-shot tool use, reversible demos, short scripts | Simple ReAct-style loop | | Branching, retries, pause/resume, audit trails, multi-agent handoffs, long-running side effects | LangGraph |

LangGraph earns its complexity when durability, inspectability, and human gates matter. If the interview problem is a single tool call with no resume story, a graph is overkill—and saying so is a strength.

## Whiteboard walkthrough: refund approval

Prompt: *Support agent drafts a refund, waits for approval, then executes.*

Sketch:

- **State keys**: `ticket_id`, `customer_id`, `refund_amount`, `draft_reason`, `approval_status`, `execution_result`, `messages` (with a merging reducer). - **Nodes**: `load_ticket` → `draft_refund` → `request_approval` (calls `interrupt` with a serializable payload: amount, reason, ticket) → `execute_refund` → `notify` → END. Rejected path: conditional edge back to `draft_refund` or to a `cancel` node. - **Checkpointer** + stable `thread_id` so the human can approve hours later. - On resume: `Command(resume=approved|rejected|edited)`. Because the approval node restarts from the top, keep the actual payout **after** the interrupt, and make any pre-interrupt logging idempotent.

That story covers state, interrupt, checkpoint, and idempotency in one diagram.

## Common failure modes

Name these without being prompted:

- `InMemorySaver` in production HITL - Non-idempotent side effects before `interrupt` - Bare `try/except` around `interrupt` - Reordering or conditionally skipping interrupts in one node - Mixing `add_edge` and `Command(goto)` from the same node - Returning `[]` to clear a merging reducer list - Cycles with no terminate condition / ignoring `recursion_limit` - Passing the wrong `Command(update=...)` shape as invoke input - Confusing LangGraph (orchestration runtime) with MCP (tooling protocol)

## What strong candidates say next

After the core design, deepen one notch: observability via LangSmith-style tracing of graph runs; **subgraphs** for multi-agent boundaries; **time travel** / `update_state` for repairing a thread; and graph migration caveats when checkpoint schemas change. Stay high level unless they pull you into implementation.

## Takeaway

Design LangGraph agents as durable state machines—shared state with reducers, explicit routing and cycles, checkpointed threads, and idempotent human gates—and reach for a simple loop when durability and pause-resume are not the problem; leave MCP to the tooling layer, not the orchestration story.

LangGraphAgentsSystem DesignHuman-in-the-LoopCheckpointsState MachinesLLM AgentsInterview Guide

Questions in this guide

Deep explanations with architecture diagrams for every question below.