튜토리얼
OpenAI Assistants API 2026-08-26 Sunset: Responses API 이전
엔드포인트 이름 변경이 아닙니다. Assistants는 4개 서버 객체, Responses는 create-run-poll-retrieve를 responses.create 한 번으로. 8월 26일 이후 유예 없이 하드 오류.
2025-08-26 OpenAI가 Assistants API beta를 2026-08-26에 영구 종료한다고 발표했습니다. assistants=v2, /v1/assistants, /v1/threads, runs는 모두 실패하며 Responses로 자동 전환되지 않습니다. Responses API + Conversations API로의 객체 매핑과 Thread 백필, JSON 출력을 정리합니다.
What actually shuts down on August 26
The entire Assistants platform goes away, not a single model. Assistant configs, accumulated Thread messages, Runs and steps—all become unreachable after sunset. Chat Completions is out of scope for this shutdown; so is the Realtime API. The trap: your app still imports OpenAI and still calls gpt-4o, but any path through beta.threads or beta.assistants dies on August 27.
OpenAI provides no automated Thread → Conversation migration. Docs state you must iterate existing Thread messages and write Conversations items yourself. Vector stores and uploaded files remain—file_search in Responses still references vs_ IDs without re-uploading.
Two dates matter: 2025-08-26 announcement plus one-year runway; 2026-08-26 hard off. New models such as GPT-5 ship only on Responses API—staying on Assistants is both a protocol and a model-access risk.
Object mapping: Assistants → Responses
The migration guide maps concepts in a table. Names align on paper; responsibilities shift:
| Assistants API | Responses era | What really changes |
|---|---|---|
| Assistant (instructions + tools) | Prompt object or inline instructions | Inline per request is safer long-term; Dashboard Prompts have their own 2026-11-30 deprecation |
| Thread (message list) | Conversation (item collection) | Items are heterogeneous—messages, function_call, tool output—not Messages alone |
| Run + poll run.status | Response (responses.create) | Sync or stream; delete the Run polling while loop |
| Run steps | Response output items | One Response bundles message, tool_call, etc. |
Assistants hosted both config and session on OpenAI—many teams never modeled thread_id as a first-class asset. Conversations still host state, but you must create conversation_id explicitly, plan archival, and Thread data will not appear by itself. If compliance needs a 2024 support Thread, export or backfill before sunset.
See OpenAI docs for details:Migrate to the Responses API and Conversations API。
Mental model: from Run polling to one Response
Classic Assistants: create thread → add message → create run → while status in queued/in_progress → list messages. Many round trips; failures attach to run_id and steps. Responses API folds “this model call” into one API: input, instructions, tools, previous_response_id or conversation binding in one request.
State does not vanish—it moves from implicit Thread to explicit Conversation or previous_response_id chains. Multi-turn agents still pass tool outputs as input items; they just stop creating Run objects. With streaming, read response.output events instead of polling.
For JSON pipelines, Structured Output is first-class on Responses: json_schema under text.format, same rules as Chat Completions response_format, different field names. OurAI Structured Output 튜토리얼covers json_schema + strict for Chat Completions; move the same schema to text.format on Responses—local validation stays the same.
Migration checklist and Thread history
Suggested order to avoid gaps:
- Grep the repo: beta.threads, beta.assistants, /v1/threads, /v1/assistants, OpenAI-Beta.*assistants, run_steps, create_and_poll
- Inventory assistant_id and thread_id in DB and env; mark active vs archivable threads
- Before sunset, export must-keep threads: list messages → Conversations items → conversations.create + items.create
- Move instructions and tools from Assistant objects to code, config, or your prompt store—not long-lived Dashboard Prompts (2026-11-30 deprecation)
- Staging: one responses.create with file_search or function tools; compare to old Assistants; validate JSON in JSONVue
- Remove run polling and run_id persistence; store response.id or conversation.id instead
Thread backfill is not copy-paste: Threads store role/content messages; Conversation items must include tool_call, tool_result types. Text-only history maps easily; code_interpreter or file_search history needs official item schemas.
Code comparison: old Assistants vs new Responses
Legacy Assistants (errors after August 26) looks like:
thread = client.beta.threads.create()
client.beta.threads.messages.create(
thread_id=thread.id,
role="user",
content="Summarize ticket #8842 for billing.",
)
run = client.beta.threads.runs.create_and_poll(
thread_id=thread.id,
assistant_id="asst_abc123",
)
messages = client.beta.threads.messages.list(thread_id=thread.id)
Equivalent Responses inlines instructions and tools, using previous_response_id or conversation for multi-turn:
response = client.responses.create(
model="gpt-4o",
instructions="You are a billing assistant.",
input="Summarize ticket #8842 for billing.",
tools=[{"type": "file_search", "vector_store_ids": ["vs_abc123"]}],
# previous_response_id="resp_..." # or conversation="conv_..."
)
print(response.output_text)
Function tools still return structured items; when arguments are JSON strings, parse and schema-validate before execution—same class of problem asStateless MCPtools/call parameter validation.
Tools, file search, and Structured Output
Vector store IDs (vs_…) continue under Responses file_search—no re-embedding. Files API uploads are shared with the Assistants era. code_interpreter is a built-in tool on Responses with different wiring—migrate tool-by-tool against current docs.
For fixed JSON shape in the final answer, use Structured Output on Responses instead of duplicating field lists in instructions. Example:
response = client.responses.create(
model="gpt-4o-2024-08-06",
input=[{"role": "user", "content": "Extract ticket fields."}],
text={
"format": {
"type": "json_schema",
"name": "support_ticket",
"strict": True,
"schema": {
"type": "object",
"properties": {
"summary": {"type": "string"},
"category": {
"type": "string",
"enum": ["billing", "bug", "other"],
},
},
"required": ["summary", "category"],
"additionalProperties": False,
},
}
},
)
Read text from output items; when the payload is JSON, JSON.parse then local schema validation. During migration, dual-run the same prompt on Assistants (pre-sunset) and Responses (staging), thenJSON Difffor key drift.
관련 글: AI Structured Output 튜토리얼, DeepSeek V4-Flash JSON 출력, Stateless MCP 해설.
FAQ
Can I still use the Assistants API after August 26?
No. All Assistants endpoints error with no degraded mode. Migrate or freeze features before sunset.
Do Threads auto-become Conversations?
No automated migration. Export Thread messages (and tool history), then write Conversations items into a new Conversation.
Is Chat Completions shutting down too?
Not in this Assistants sunset. OpenAI steers new features to Responses, but Chat Completions integrations need not flip the same day unless you need Responses-only models or tools.
Are Prompt objects the same as Assistants?
Not quite. Prompts are versioned dashboard configs with their own deprecation timeline. Safer for long-lived code: keep instructions and tools in your config and pass them on each responses.create.
Summary and next steps
Assistants sunset is an architecture swap: four server objects become Responses + Conversations, Run polling disappears, Thread history does not follow automatically. Vector stores reuse; JSON Schema contracts port. Real work is finding thread/assistant dependencies, backfilling history, and rewriting orchestration.
Next: grep for stragglers, responses.create in staging, paste model JSON into JSONVue. For schema constraints read Structured Output; for remote tool protocol read Stateless MCP.