チュートリアル
DeepSeek V4-Flash:API・Agent・Tool Calling・JSON 出力の変更
V4-Flash は名称変更だけではありません。モデル ID、Thinking + Tool Calling の再生ルール、JSON モード(json_schema なし)が変わります。Tool Calls と JSON 出力の使い分け、Agent ループで必須のフィールドが分かります。
2026 年 DeepSeek は API 主軸を V4 系へ:deepseek-v4-flash と deepseek-v4-pro が deepseek-chat / deepseek-reasoner に代わります。Chat Completions は OpenAI 互換のままですが、Thinking では reasoning_content の完全再生が必要。Tool Calling は標準 tools、strict(Beta)で引数 Schema。JSON は response_format: json_object のみ。Structured Output や MCP の文脈と合わせて API・Agent・JSON の三層を整理します。
V4-Flash 更新の要点
公式 Changelog:呼び出し方法は同じで model を deepseek-v4-flash に。V4-Flash-0731 は Preview と同アーキテクチャの再 post-train。Responses API 形式をネイティブサポート。deepseek-chat / deepseek-reasoner は 2026-07-24 廃止予定。Thinking + Tools の再生ロジックは必ず再テストを。
JSON 形式精度は内測で 78%→85%、正規表現で約 97%。IFEval Prompt-Level は 63.9%→77.6%。Schema 制約ではなく、json_object は構文 JSON の保証に留まります。
コンテキスト 1M、最大出力 384K。JSON モードでは max_tokens 不足で JSON が途中切断されるため、タスクに応じた上限設定が必要です。
API 接続:モデル名と互換性
base_url は https://api.deepseek.com、/chat/completions 不変。OpenAI SDK で base_url と model を変更。GET /models は deepseek-v4-flash / deepseek-v4-pro が中心。旧エイリアスに依存しないでください。
最小の非 Thinking テキスト完了リクエストの骨格です。
from openai import OpenAI
client = OpenAI(
api_key="YOUR_KEY",
base_url="https://api.deepseek.com",
)
resp = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "user", "content": "Summarize JSON tooling in one sentence."}
],
max_tokens=256,
)
print(resp.choices[0].message.content)
LangChain 等では thinking フィールドのマッピングと reasoning_content 破棄に注意。2 ターン目で HTTP 400 になることがあります。
公式 Changelog と Tool Calls ガイド: DeepSeek API Updates と Tool Calls.
Thinking モードと Agent ループ
extra_body={"thinking": {"type": "enabled"}} で Thinking 有効。content に加え reasoning_content が返ります。
tool_calls 付き assistant メッセージでは、以降の全リクエストで reasoning_content を完全に再生。欠落すると HTTP 400。
role、content、reasoning_content、tool_calls をセットで保存し、API 順序どおり replay。推論フィールドを削らない。
tool_calls が null/空なら終了。orchestrator は content を最終回答に、または JSON 抽出へ。
Tool Calling:引数 JSON と strict
OpenAI 標準の tools。function.arguments は JSON 文字列—parse と検証は実行側の責務。
Function Calling strict(Beta)は引数 Schema に沿わせます。Beta のため Schema サブセット制限あり。
モデルは関数を実行しません。MCP tools/call と同様、引数 JSON はローカル Schema 検証を。
外部データは Tool Calls、コンテキスト抽出は json_object。Agent は中間 Tool、最終 json_object が定番。
JSON 出力:json_object と空レスポンス
response_format: {"type": "json_object"} のみ。prompt に json と例、max_tokens 設定、空 content 時のリトライが必須。
Thinking 併用可。空 content は明示的 system 指示で改善。parse 後も Schema 検証を。
下記:チケット意図と緊急度の抽出例(非 Thinking)。
resp = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{
"role": "system",
"content": (
'Extract intent and urgency as json. '
'Example: {"intent":"refund","urgency":"high"}'
),
},
{"role": "user", "content": "Card charged twice, need refund today."},
],
response_format={"type": "json_object"},
max_tokens=400,
)
OpenAI/Gemini Structured Output との差は AI Structured Output チュートリアル参照。DeepSeek は構文 JSON 層まで。
JSON 出力の公式ドキュメント: JSON Output.
Structured Output との境界
ベンダー比較表:
| シナリオ | DeepSeek V4-Flash | OpenAI / Gemini Structured Output |
|---|---|---|
| 最終回答の固定フィールド | json_object + 例 + ローカル Schema | json_schema + strict、デコード時形状固定 |
| Agent が外部ツール呼び出し | tools + reasoning_content 再生 | 同様 tools、一部 strict 対応 |
| Remote MCP / JSON-RPC | 公式 MCP なし、HTTP は自作 | クライアント側 MCP、Stateless MCP 記事参照 |
MCP + Tool Calls + 最終 json_object の三層構成も可能。いずれも JSON 検証を。
関連記事:AI Structured Output チュートリアル、Gemini API JSON 出力、Stateless MCP 解説。
よくある質問 FAQ
deepseek-chat はまだ使えますか?
2026-07-24 廃止予定。deepseek-v4-flash / v4-pro へ移行し、Thinking + Tools を staging で再テスト。
Tool Calls と JSON を同時に?
分離設計が無難:ループは tools、最終構造化は json_object。
JSON モード後も Schema 検証が必要な理由
構文 JSON のみ保証。キー名・型・enum は保証されません。
reasoning_content の保持期間
tools 継続可能なスレッドでは完全 replay が必要。アクティブ Agent からは削除しない。
まとめと次の一歩
V4-Flash は OpenAI 互換のまま Agent 向け契約が厳格化。reasoning_content 再生、arguments 文字列、空/切断 JSON に注意。
次:Tool チェーンと JSON 抽出を試し JSONVue で検証。Structured Output / Gemini / Stateless MCP 記事も参照。