教程

DeepSeek V4-Flash 最新消息:API、Agent、Tool Calling 和 JSON 輸出有什麼變化?

V4-Flash 不是換皮:模型 ID 換了,Thinking 與 Tool Calling 的回放規則變了,JSON 模式仍只有 json_object 沒有 Schema 約束。讀完你能判斷該走 Tool Calls 還是 JSON 輸出,以及 Agent 循環裏哪些字段必須原樣回傳。

DeepSeek 在 2026 年把 API 主線切到 V4 系列:deepseek-v4-flash 與 deepseek-v4-pro 取代即將停用的 deepseek-chat 和 deepseek-reasoner。對工程團隊來說,變化不在「多一個模型名」,而在三條鏈路的契約——Chat Completions 仍 OpenAI 兼容,但 Thinking 模式要求完整回放 reasoning_content;Tool Calling 走標準 tools 參數,strict 模式(Beta)才約束參數 Schema;JSON 輸出只有 response_format: json_object,沒有 OpenAI 那套 json_schema。如果你已經在用 DeepSeek 做抽取、分類或 Agent,這篇把 API 面、Agent 面、JSON 面拆開講,並接上站內 Structured Output 與 MCP 的上下文。

V4-Flash 更新在改什麼

官方 Changelog 把這次更新概括成幾條硬事實:調用方式不變,把 model 設爲 deepseek-v4-flash 即可;V4-Flash-0731 與 Preview 同架構同尺寸,只是重新 post-train;V4-Flash 原生支持 Responses API 格式並針對 Codex 做了適配;deepseek-chat 與 deepseek-reasoner 計劃在 2026-07-24 停用。這意味着現有集成大多隻需改 model 字符串,但 Thinking + Tools 的組合必須重新測一遍回放邏輯。

性能側,文檔提到 JSON 格式輸出準確率在內測集從 78% 提到 85%,配合正則可進一步到 97%;IFEval Prompt-Level 從 63.9% 跳到 77.6%。這些數字說明 V4-Flash 更聽 system 指令,但不等於 Schema 級約束——JSON 模式仍只保證「能 parse」,不保證字段名和類型。

上下文窗口文檔寫 1M token 輸入、最大輸出 384K token。生產裏仍應按任務設 max_tokens,尤其在 JSON 模式下,官方明確警告:token 不夠會把 JSON 字符串截斷在中間。

API 接入:模型名與兼容層

base_url 仍是 https://api.deepseek.com,Chat Completions 端點 /chat/completions 不變。OpenAI SDK 只需改 base_url 和 model;Anthropic 兼容接口也可訪問 V4 模型。GET /models 列出的當前模型通常是 deepseek-v4-flash 與 deepseek-v4-pro,不要繼續依賴舊別名。

下面是最小請求骨架:非 Thinking、純文本.completion。

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KEY",
    base_url="https://api.deepseek.com",
)

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "user", "content": "Summarize JSON tooling in one sentence."}
    ],
    max_tokens=256,
)
print(resp.choices[0].message.content)

若你用的是 LangChain、LiteLLM 或自研網關,先確認框架有沒有把 DeepSeek 的 thinking 字段映射錯——有些配置會禁用 tool_choice 或丟棄 reasoning_content,導致 Agent 第二輪直接 400。

官方 Changelog 與 Tool Calls 指南見 DeepSeek API Updates 與 Tool Calls.

Thinking 模式與 Agent 循環

V4-Flash 與 V4-Pro 都支持 Thinking 與非 Thinking。開啓方式是在請求里加 extra_body,例如 extra_body={"thinking": {"type": "enabled"}}。響應裏除了 content,還會出現 reasoning_content——模型內部推理鏈,不一定是你想展示給用戶的文案。

Agent 循環的關鍵規則:只要某輪 assistant 消息帶 tool_calls,後續所有請求必須把該 assistant 消息的 reasoning_content 完整回傳。漏傳會 HTTP 400。這和 OpenAI 近期「保留 reasoning items」的思路類似,但 DeepSeek 在帶 tools 參數時寫得更硬。

實踐上,你的會話存儲結構要能存三元組:role、content、reasoning_content、tool_calls。Replay 時按 API 要求的順序 append,不要自己 strip 推理字段「省 token」——省下來的字節會在下一輪變成 400。

終止條件:當 assistant 消息的 tool_calls 爲 null 或空,模型認爲任務完成,不再請求外部工具。你的 orchestrator 應在此刻讀取 content 做最終答覆或再開 JSON 抽取。

Tool Calling:參數 JSON 與 strict

Tool Calling 走 OpenAI 標準:請求裏聲明 tools 數組,每項含 type: function、name、description、parameters(JSON Schema)。模型返回 assistant.tool_calls,其中 function.arguments 是 JSON 字符串,不是已 parse 的對象——你的執行層必須 JSON.parse 並校驗後再調真實 API。

DeepSeek 文檔提供 Function Calling strict 模式(Beta):在工具定義上打開 strict,服務端會盡量讓 arguments 貼合聲明的 Schema。Beta 意味着 Schema 子集受限,不是任意 JSON Schema 都能過——集成測試應 pinned 到文檔列出的支持關鍵詞。

記住:模型不執行函數。它只決定「要不要調、調哪個、參數 JSON 長什麼樣」。查庫存、寫數據庫、發 HTTP 都是你的後端職責;arguments 落地前仍建議走 JSON Schema 校驗,和 MCP tools/call 的參數校驗是同一類工程問題。

和 JSON 輸出分工:需要實時外部數據時用 Tool Calls;只需要從已有上下文抽字段時用 json_object。Agent 常見模式是中間步 Tool Calls,最後一步 json_object 輸出結構化報告。

JSON 輸出:json_object 與空響應

DeepSeek JSON 模式只有 response_format: {"type": "json_object"},沒有 response_format: json_schema。要穩定拿到 JSON,官方四條要求都要滿足:設置 response_format;在 system 或 user 裏出現「json」字樣並給出示例形狀;合理設置 max_tokens 防截斷;知道 API 偶爾返回空 content,生產要有重試或降級。

Thinking 與 JSON 可以同時開,但 json_object 不能替代清晰指令。實測裏出現過 HTTP 200、finish_reason: stop、content 爲空的情況;加強 system 裏「返回一個合法 JSON 對象,鍵爲 …」通常能修復。解析後仍要本地校驗鍵和類型——JSON 模式不保證 Schema。

下面示例:抽取工單意圖與緊急程度,非 Thinking,帶寬 max_tokens。

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {
            "role": "system",
            "content": (
                'Extract intent and urgency as json. '
                'Example: {"intent":"refund","urgency":"high"}'
            ),
        },
        {"role": "user", "content": "Card charged twice, need refund today."},
    ],
    response_format={"type": "json_object"},
    max_tokens=400,
)

拿到字符串後,先 JSON.parse,再丟進 Schema 校驗器。和 Gemini、OpenAI Structured Output 的差別,見站內 AI Structured Output 教程——DeepSeek 目前停在「語法 JSON」層,形狀約束要靠 prompt 示例 + 本地校驗,或 Tool strict(僅參數側)。

JSON 輸出官方說明見 JSON Output.

和 Structured Output 的邊界

可以把各家能力放在同一張表裏對比選型:

場景 DeepSeek V4-Flash OpenAI / Gemini Structured Output
最終答覆要固定字段 json_object + prompt 示例 + 本地 Schema 校驗 json_schema + strict,解碼階段約束形狀
Agent 調外部工具 tools + 回放 reasoning_content 同類 tools 參數;部分棧支持參數 Schema strict
Remote MCP / JSON-RPC 無官方 MCP;HTTP 工具需自建 客戶端側 MCP;協議 JSON 見 Stateless MCP 文

三條鏈路可以並存:MCP Server 暴露工具,DeepSeek Agent 用 Tool Calls 調你的 HTTP 適配層,最後一跳 json_object 彙總成報表 JSON。無論哪一層,JSON 字符串落地前都建議在瀏覽器或 CI 裏做一次校驗。

延伸閱讀:AI Structured Output 教程、Gemini API JSON 輸出、Stateless MCP 解析。

常見問題 FAQ

deepseek-chat 還能用嗎?

官方計劃 2026-07-24 停用 deepseek-chat 與 deepseek-reasoner。新集成請直接用 deepseek-v4-flash 或 deepseek-v4-pro,並在 staging 跑一遍 Thinking + Tools 回放測試。

Tool Calls 和 JSON 模式能同請求開嗎?

通常分開設計:Agent 循環用 tools,最終結構化彙總用 json_object。同請求既 tools 又 json_object 會讓調試變複雜,且 Thinking 模式下回放字段更多。

爲什麼 JSON 模式仍要 Schema 校驗?

json_object 只保證合法 JSON,不保證鍵名、類型、枚舉。模型可能給出語義對但字段漂移的對象。本地 Schema 校驗是 cheap 的保險,換模型或關 strict 時尤其有用。

reasoning_content 要存多久?

只要會話還可能繼續帶 tools 調用,就要能完整 replay 歷史 assistant 消息裏的 reasoning_content。會話歸檔策略可以壓縮舊輪次,但活躍 Agent 線程不要丟該字段。

總結與下一步

DeepSeek V4-Flash 把模型能力推到 V4 命名空間,API 外殼仍 OpenAI 兼容,但 Agent 開發者必須處理 reasoning_content 回放、Tool arguments 的 JSON 字符串解析,以及 json_object 的空響應與截斷風險。JSON 準確率提升不等於 Schema 約束——要和 OpenAI/Gemini Structured Output 區分層級。

下一步:用 deepseek-v4-flash 跑一條 Tool Call 鏈和一條 JSON 抽取,把返回值貼進 JSONVue 做格式化和 Schema 校驗。需要對比解碼級 Schema 約束時,閱讀 AI Structured Output 與 Gemini JSON 輸出教程;需要 Remote 工具部署時,閱讀 Stateless MCP 文。