Tutorial

DeepSeek V4-Flash: API, Agent, Tool Calling, JSON

V4-Flash ist mehr als ein Rename: neue Modell-IDs, Replay-Regeln für Thinking + Tools, JSON-Modus nur json_object ohne Schema-Zwang. Tool Calls vs. JSON-Ausgabe und Pflichtfelder in der Agent-Schleife werden klar.

2026 verlagert DeepSeek die API auf V4: deepseek-v4-flash und deepseek-v4-pro ersetzen deepseek-chat und deepseek-reasoner. Chat Completions bleibt OpenAI-kompatibel, Thinking verlangt vollständiges reasoning_content-Replay, Tool Calling nutzt tools mit strict (Beta), JSON nur json_object. API-, Agent- und JSON-Ebene im Kontext von Structured Output und Stateless MCP.

What the V4-Flash update changes

The official changelog states hard facts: calling style is unchanged—set model to deepseek-v4-flash; V4-Flash-0731 matches Preview architecture and size with a re-post-train; V4-Flash natively supports the Responses API shape with Codex-oriented config; deepseek-chat and deepseek-reasoner are scheduled to retire on 2026-07-24. Most integrations only change the model string, but Thinking + Tools combinations need a fresh replay test.

On metrics, docs cite JSON format accuracy rising from 78% to 85% on an internal set, and up to ~97% with regex cleanup; IFEval Prompt-Level jumps from 63.9% to 77.6%. That means V4-Flash follows system instructions better—not schema-level guarantees. JSON mode still means “parsable JSON”, not field names and types.

Docs list a 1M-token context and up to 384K output tokens. Production should still set task-appropriate max_tokens; in JSON mode the API warns that too small a budget truncates the JSON string mid-flight.

API access: model names and compatibility

base_url remains https://api.deepseek.com and Chat Completions stay at /chat/completions. Point the OpenAI SDK at that base_url and model; Anthropic-compatible endpoints also expose V4. GET /models typically lists deepseek-v4-flash and deepseek-v4-pro—do not rely on legacy aliases.

Below is a minimal non-Thinking text completion skeleton.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KEY",
    base_url="https://api.deepseek.com",
)

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "user", "content": "Summarize JSON tooling in one sentence."}
    ],
    max_tokens=256,
)
print(resp.choices[0].message.content)

If you use LangChain, LiteLLM, or a home gateway, verify the framework maps DeepSeek’s thinking fields correctly—some configs disable tool_choice or drop reasoning_content, yielding HTTP 400 on turn two of an agent.

See the official changelog and Tool Calls guide: DeepSeek API Updates and Tool Calls.

Thinking mode and the agent loop

V4-Flash and V4-Pro support Thinking and non-Thinking. Enable via extra_body, e.g. extra_body={"thinking": {"type": "enabled"}}. Besides content, responses include reasoning_content—the internal chain, not necessarily user-facing copy.

Agent rule: whenever an assistant message carries tool_calls, every later request must replay that message’s reasoning_content intact. Omit it → HTTP 400. Similar in spirit to OpenAI’s “keep reasoning items”, but DeepSeek is explicit when tools is present.

Store role, content, reasoning_content, and tool_calls together. Replay in API order; do not strip reasoning to “save tokens”—you pay with 400 on the next turn.

Stop when assistant.tool_calls is null or empty—the model considers tools done. Your orchestrator reads content for the final answer or opens a JSON extraction pass.

Tool calling: argument JSON and strict mode

Tool calling follows OpenAI: declare tools with type function, name, description, parameters (JSON Schema). The model returns assistant.tool_calls where function.arguments is a JSON string—not a parsed object. Parse and validate before hitting real APIs.

DeepSeek documents Function Calling strict mode (beta): strict on the tool definition pushes arguments toward the declared schema. Beta means a limited schema subset—not every JSON Schema keyword passes—pin integration tests to documented support.

The model never executes functions. It only chooses whether to call, which tool, and the argument JSON shape. Inventory, databases, and HTTP are your backend; validate arguments with JSON Schema just like MCP tools/call parameters.

Split concerns: live external data → Tool Calls; field extraction from context → json_object. Common agent pattern: Tool Calls mid-flight, json_object for the final structured report.

JSON output: json_object and empty responses

DeepSeek JSON mode is only response_format: {"type": "json_object"}—no json_schema variant. Official guidance: set response_format; include the word “json” plus an example shape in system or user; set max_tokens to avoid truncation; expect occasional empty content and retry or degrade in production.

Thinking and JSON can combine, but json_object is not a substitute for a clear instruction. Live tests saw HTTP 200, finish_reason stop, and empty content; an explicit system line like “return one valid JSON object with keys …” usually fixes it. After parse, validate keys and types locally—JSON mode is not a schema.

Example below: extract ticket intent and urgency, non-Thinking, generous max_tokens.

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {
            "role": "system",
            "content": (
                'Extract intent and urgency as json. '
                'Example: {"intent":"refund","urgency":"high"}'
            ),
        },
        {"role": "user", "content": "Card charged twice, need refund today."},
    ],
    response_format={"type": "json_object"},
    max_tokens=400,
)

After you have a string, JSON.parse then run a schema validator. For decode-level schema constraints vs Gemini/OpenAI Structured Output, see our AI Structured Output guide—DeepSeek today stops at syntactic JSON; shape control is prompt examples + local validation, or Tool strict on parameters only.

Official JSON output docs: JSON Output.

Boundary vs Structured Output

Compare vendors on one table:

Scenario DeepSeek V4-Flash OpenAI / Gemini Structured Output
Fixed fields in the final answer json_object + prompt example + local Schema check json_schema + strict, shape at decode time
Agent calls external tools tools + replay reasoning_content Similar tools; some stacks strict on parameter schemas
Remote MCP / JSON-RPC No official MCP; HTTP tools are DIY Client-side MCP; protocol JSON in our Stateless MCP article

All three can coexist: an MCP server exposes tools, a DeepSeek agent calls your HTTP adapter via Tool Calls, and a final json_object hop emits the report JSON. Whichever layer, validate JSON strings before production.

Further reading: AI Structured Output tutorial, Gemini API JSON output, and Stateless MCP guide.

FAQ

Can I still use deepseek-chat?

Official plan retires deepseek-chat and deepseek-reasoner on 2026-07-24. New work should use deepseek-v4-flash or deepseek-v4-pro and rerun Thinking + Tools replay tests in staging.

Tool Calls and JSON mode in one request?

Design them separately: tools for the agent loop, json_object for the final structured summary. Both in one request complicates debugging and adds replay fields under Thinking.

Why validate Schema after JSON mode?

json_object guarantees valid JSON, not key names, types, or enums. The model may drift fields while staying syntactic. Local schema validation is cheap insurance when models or strict flags change.

How long must I keep reasoning_content?

For any thread that may continue with tools, you must replay full reasoning_content on historical assistant messages. Archive policies may compress old turns, but active agent threads must not drop the field.

Summary and next steps

DeepSeek V4-Flash moves capability under V4 names while keeping an OpenAI-compatible shell. Agent builders must handle reasoning_content replay, JSON-string tool arguments, and json_object empty/truncation risks. Higher JSON accuracy is not schema constraint—keep layers separate from OpenAI/Gemini Structured Output.

Next: run one Tool Call chain and one JSON extraction with deepseek-v4-flash, paste results into JSONVue for format and schema checks. For decode-level schema constraints read AI Structured Output and Gemini JSON guides; for remote tools read Stateless MCP.