Tutorial
What is AI Agent State? 2026 guide to state management, JSON State, memory, workflows, and task execution
Chat history is not state. An agent can pause, resume, and hand off only if a checkable JSON snapshot exists: goal, step, tool receipts, memory pointers, and a task envelope.
Shipping agents in 2026 rarely fails because the model is “not smart enough.” It fails after a crash, a retry, or a human approval — when the runtime no longer knows which hop it was on. The previous guide folded an agent into an observe–decide–call loop; see What is an AI Agent. This article adds the layer outside the loop: state. State is not the whole transcript stuffed into context, and it is not a vector snippet. It is a structured snapshot of the runtime’s current position — almost always JSON. We split JSON State, memory, workflows, and task execution, then link our Stateless MCP, A2A, Structured Output, and 1M-context guides.
What agent state is — vs chat history and sessions
One sentence is enough: agent state is the structured snapshot that makes a run pausable, resumable, and replayable. It answers four questions: what is the goal, which hop are we on, which work results are already committed, and who is blocking the next step (a tool, a human, or maxSteps). The model chooses an action from the current observation; the runtime executes and writes the position back. Without a checkable position, the loop can only paste the whole chat back in — that is chat, not state management.
Chat history is the message list the model sees: user / assistant / tool. A session (OpenAI Conversations, previous_response_id, the old Assistants Thread) is the handle the vendor sees: which model-side items the next request should carry. State is the position your runtime sees: how far reconciliation got, the amount awaiting confirm, the current graph node, the checkpoint id. They can coexist; they must not impersonate each other. Treat messages[] as the only source of truth and retries double-charge or double-email. Treat a Conversation id as business state and a vendor swap drops the position. After Assistants moved to Responses, the session primitive changed — bind business snapshots even less to platform objects. See Assistants → Responses migration.
Most 2026 incidents are mixed layers: tool arguments written into state, the whole state blob stuffed into the next prompt, or long-term memory replayed as a checkpoint. Once you split them, debugging has a handle: the model misread an observation; the reducer / Schema wrote a bad position; memory recalled the wrong preference. The map:
| Concept | Who reads it | If it is lost |
|---|---|---|
| Chat / messages | The model (this turn’s context) | Wrong answers; usually re-projectable from a checkpoint |
| Session / Conversation | The model vendor | Reasoning items drift; your business position should still stand alone |
| Agent state / checkpoint | Your runtime and orchestrator | Unsafe resume; retries may double-write downstream |
| Memory | Cross-run retrieval and preferences | Forgotten habits; never the truth of the current step |
JSON State: the checkpoint is the contract
You write state as JSON so you can validate, diff, and replay. Orchestrators in the LangGraph family persist a checkpoint every super-step, thread them with thread_id, and point at one frame with checkpoint_id. Production uses a Postgres or SQLite saver, not an in-process MemorySaver. Names change; the shape should not: a parseable object with schemaVersion, a status enum, step, working, memoryRefs. The serializer may be JsonPlus or extended JSON, but the layer you log, expose, and debug should be a plain JSON object — otherwise Schema checks and a browser Diff are useless.
The snapshot below is position only. It does not embed messages[] or raw downstream HTTP. Business fields live in working; a tool call keeps name and callId; memory is pointers. Project the full transcript from the checkpoint into model context when needed — do not treat context as state in reverse.
{
"schemaVersion": "1.0",
"runId": "run_7c2a",
"threadId": "thr_invoice_42",
"goal": "Reconcile January 2026 paid invoices",
"status": "awaiting_tool",
"step": 3,
"maxSteps": 12,
"node": "call_tools",
"plan": ["searchInvoices", "sumTotals", "askConfirm"],
"working": {
"invoiceCount": 2,
"currency": "USD"
},
"pendingTool": {
"name": "searchInvoices",
"callId": "call_8f3a"
},
"memoryRefs": ["mem_user_prefs", "mem_last_reconcile"],
"checkpointId": "ckpt_3"
}
Both styles work: a full snapshot each time, or a JSON Patch / reducer. Full snapshots diff and replay cleanly; patches save storage and must themselves be validatable. Either way, lock a Draft 2020-12 Schema: status as an enum (running / awaiting_tool / awaiting_human / succeeded / failed / cancelled), step as an integer, working keys from the business, additionalProperties false so tool junk cannot leak in. The answer to the user or a downstream system is a separate Structured Output file — do not share it with the checkpoint. See AI Structured Output.
Memory: recall is not the machine position
Memory answers “what should still be known across time.” State answers “where this frame stopped.” Three layers are common in 2026: working memory (this turn’s messages and recent tool_result), short-term / thread memory (the checkpoint chain on one thread_id — LangGraph’s short-term memory), and long-term memory (a cross-thread Store: preferences, facts, procedures). Stuffing all three into one giant JSON looks cheap; replay and forgetting policies break together.
By content: episodic (what happened in this reconciliation), semantic (the user prefers USD totals), procedural (reusable steps such as “search invoices, then sum”). Only memories cited by the current run belong in State.memoryRefs. A million-token window does not retire those pointers — a window is a budget, not the source of truth. See 1M token context.
| Layer | Typical carrier | Do not put here |
|---|---|---|
| Working memory | This turn’s messages / tool_result | Full cross-session user preferences |
| Short-term / checkpoint | State snapshots on a thread_id | Raw vector-store text, untrimmed logs |
| Long-term store | Items keyed by userId / namespace | Current step, pendingTool, idempotency keys |
| Vendor session | Conversation / previous_response_id | Your business working object |
Give long-term items a stable envelope: id, kind, scope, text or data, source, updatedAt. source says whether the user stated it, a tool returned it, or the model summarized it — the last two must be revocable. After a retrieval hit, write only the id into State and inject the body into context on demand. Example:
{
"id": "mem_user_prefs",
"kind": "semantic",
"scope": "user",
"userId": "u_1042",
"text": "Prefers USD totals and weekday email summaries",
"source": "explicit_setting",
"updatedAt": "2026-09-01T09:00:00Z"
}
Workflow vs agent: who draws edges, who writes state
In a classic workflow (n8n, Temporal, a homegrown state machine) people pre-wire the next hop. State is workflow variables: order id, retry count, whether compensation fired. In an agent, the model chooses the next hop from the current JSON observation, so state also records plan, node, pendingTool. Production in 2026 is rarely purebred: a graph with agent nodes — edges are workflow, the inside of a node is a tool loop. One JSON state object serves two readers: the orchestrator reads status / node; the model sees only the observation subset you project.
MCP does not own this. Remote MCP around 2026-07-28 is stateless JSON-RPC: each request carries its own metadata; the server does not remember your business position. That is the right transport choice, not “agents cannot have state.” Application state still lives in your checkpoint. Details: Stateless MCP. Tool-argument Schema and state Schema must be separate files — one hop’s inputs vs the machine position. See MCP and JSON Schema.
Sideways delegation uses A2A: the peer is an opaque agent; the task has its own lifecycle (submitted / working / completed / failed) and artifacts. That is another state machine — do not merge it into the local checkpoint. The orchestrator pins them with parentRunId. Compare A2A vs MCP. A model gateway (a local /v1, for example) only changes inference supply, not state shape. Stable contracts make fallback meaningful.
Task execution: runs, steps, idempotency, retries
The execution layer turns a snapshot into a recoverable machine. Each user goal opens a runId; each tool hop in the loop is a step; every downstream write (charge, email, ticket) carries an idempotencyKey. After a crash, resume from the latest checkpoint and do not replay side effects of steps that already succeeded. Pending writes (some nodes ok, some failed) belong in the snapshot — not in an operator’s guess.
Human approval is a first-class status, not a special branch: status=awaiting_human, the object to confirm sits in working, and resume allows only legal transitions (approve → continue, reject → fail or replan). Do not “ask the model again” instead of a state transition — the model cannot see a click you never wrote into the observation. maxSteps, user cancel, and Schema failure are stop conditions too: write them into status, not only logs.
Pick one owner for history between the vendor session and your execution record. OpenAI Responses can continue reasoning items via a Conversation or previous_response_id; that tape is model-side, not invoice-reconciliation position. Recommended: you own the checkpoint and the task envelope; the platform owns only the reasoning items it requires you to replay (some vendors demand reasoning_content verbatim whenever tool_calls are present). A task envelope:
{
"taskId": "task_a2a_91",
"parentRunId": "run_7c2a",
"kind": "delegate",
"status": "working",
"idempotencyKey": "inv-jan-2026-reconcile",
"steps": [
{ "id": "s1", "name": "searchInvoices", "ok": true },
{ "id": "s2", "name": "sumTotals", "ok": null }
],
"artifacts": []
}
When you delegate to a child agent, store the remote taskId in local working or steps — do not flatten their artifacts into the same checkpoint. Intermediate products are a new JSON family; keep them out of final Structured Output and MCP arguments. Stamp correlationId / runId on every tool log now so later reconciliation maps. Platform observability (DevDay-style dashboards) should join on the same key. See OpenAI DevDay 2026 predictions.
Validation on the ground and JSONVue
Before a checkpoint lands: parse → Schema → business rules. A clever model does not replace those three steps. A bad state is worse than bad arguments: arguments are one hop; state is the run’s source of truth.
- JSON.parse the checkpoint; on failure refuse the write and keep the previous ckpt_id.
- Validate status / step / working against the State Schema (Draft 2020-12); emit path and keyword.
- Business gate: monotonic step, stable idempotency key, every memoryRef exists, illegal status transitions fail closed.
Line up three blobs: ckpt_n, ckpt_n+1, and the observation you projected to the model. Shape jumps are almost always the reducer. In the browser: JSON formatter to read the snapshot tree; JSON Schema validator to lock State and Memory envelopes; JSON Diff for adjacent checkpoints. Share valid / missing-step / illegal-status fixtures in CI and manual debug.
Further reading: What is an AI Agent, Structured Output, Stateless MCP, A2A vs MCP, 1M token context.
FAQ
Are state and memory the same thing?
No. State is the current run’s position (can you resume safely?). Memory is recall across time (preferences, facts, old episodes). A checkpoint chain can serve as short-term memory; a long-term store must not hold step / pendingTool. Replay uses state; retrieval uses memory.
Context windows hit 1M — do we still need checkpoints?
Yes. A window decides how much observation this turn can hold. It does not decide which hop to resume after a crash, or whether a retry double-writes. Treating the whole history as state worsens both the bill and the failure surface. 1M is a budget tool; a checkpoint is an execution tool.
We use OpenAI Conversations — do we still store JSON state?
Yes. Conversation / previous_response_id continues model-side items, not your business position. After a quota, region, or gateway change, the vendor session may not line up. working, idempotency keys, and human-approval status belong on JSON you control.
Does stateless MCP mean we cannot build a stateful agent?
No. Protocol-stateless only means each tools/call carries its own arguments; the server does not store your reconciliation progress. Application state lives in your checkpoint; MCP remains discovery and transport. Keep inputSchema and the State Schema in two files.
Summary and next steps
2026 agent state in one line: the runtime remembers position in a validatable JSON snapshot, memory only supplies cited recall, a workflow draws edges or hands them to the model, and execution turns the snapshot into a recoverable machine with run / step / idempotency keys. Chat history and vendor sessions do not replace that snapshot.
Next: write your State Schema and one valid checkpoint; Diff adjacent snapshots in JSONVue; add missing-field and illegal-status fixtures. Loop definition in the Agent article; final-answer shape in Structured Output; protocols in MCP / A2A.