Guide

OpenAI DevDay 2026 developer guide: what may change in Responses, Structured Outputs, Tool Calling, and MCP

The keynote will not edit your repo. What you can do now is collapse final output, tool arguments, MCP inputSchema, and session items into one checkable contract.

OpenAI DevDay 2026 is 29 September at Fort Mason in San Francisco. The official page promises a technical day on APIs and developer tools: a livestreamed morning keynote (including Sam Altman), afternoon breakouts, and recordings afterwards. Applications closed in July. The 2 September piece listed ten changelog predictions—Completions dates, Ultrafast, enterprise identity, Realtime parity; see DevDay 2026 API predictions. This is the handbook: four lines that change the shape of your code—Responses, Structured Outputs, Tool Calling, MCP—and the schema files you can freeze this week. Predictions miss. Split contracts do not. Facts still come from the OpenAI API changelog: Assistants shut down for good on 26 August; GPT-5.6 put programmatic tool calling and multi-agent orchestration on Responses; remote MCP, Skills, computer use, and tool search almost never land on Chat Completions.

How to use this guide vs the forecast

The forecast answers “what the stage might say.” This guide answers “which JSON belongs in the repo this week.” Both follow the same 2026 trajectory; the reader action differs. Use the former against the keynote. Use the latter to freeze a schemas/ tree. Do not rewrite the product for a prediction. DevDay rarely invents a surface—it graduates previews, drops Beta labels, and folds enterprise allowlists into defaults.

Separate facts from maybes. Already true: Responses is the Agent-capable path; Structured Outputs pins the final object with text.format.json_schema + strict; function tools and remote MCP can share one responses.create; MCP asks for approval by default (mcp_approval_request) and can be narrowed with require_approval / allowed_tools. None of that is a prediction. What may be announced: a sunset date, a GA sticker, a self-serve toggle, and an official way to declare two schemas as the same origin.

The useful question is not “will GPT-5.7 appear” but: which request surface becomes the only recommended path? Which JSON Schema flips from model hint to platform constraint? Which session state moves from your database into official Conversations? The map:

Document What it answers What you do now
DevDay forecast (2 September)Odds and evidence for ten API directionsScore the keynote; do not reshuffle the repo
This guide (15 September)Four tracks + schema files to freezeSplit schemas/tools, output, mcp, session this week
Assistants migrationHow Thread / Run move onto ResponsesClear beta.threads; treat session items as a contract
Changelog / official docsWhat is GA, deprecated, or previewTrust the docs, not a social recap

Responses API: the request surface that may consolidate

Assistants is already dead. Chat Completions is alive, but 2026 features almost never land there: remote MCP, tool search, computer use, Skills, hosted shell, WebSocket Responses, and phase (commentary / final_answer) all hang off Responses. Reusable prompts never shipped on Completions. That is platform policy, not taste: the Agent-capable surface is being collapsed to one. For the object move from Assistant / Thread / Run, see Assistants → Responses migration.

The likeliest DevDay cuts are not “one more Responses parameter” but some of these three: a dated Completions freeze or shutdown; Conversations as the session primitive across text Responses and Realtime; more image, transcription, and video work as built-in tools on the same output-item timeline. For you that means new wrappers speak only input / output items, tools[], text.format, and previous_response_id or conversation. Do not keep two tools[] shapes.

Request metadata consolidates too. Fast already replaced Priority; Ultrafast is still a limited preview. Put service_tier, prompt_cache_retention, and safety_identifier in the request JSON now, not in SDK defaults. After the keynote you will match invoices, cache hits, and safety blocks by those keys—not by the model name. The skeleton below is legal today and should remain legal: function tools, remote MCP, and Structured Output on one responses.create.

{
  "model": "gpt-5.6-sol",
  "input": "Extract the paid invoice and call billing tools",
  "tools": [
    {
      "type": "function",
      "name": "searchInvoices",
      "strict": true,
      "parameters": {
        "type": "object",
        "properties": {
          "invoiceId": { "type": "string" },
          "status": {
            "type": "string",
            "enum": ["draft", "sent", "paid", "void"]
          }
        },
        "required": ["invoiceId", "status"],
        "additionalProperties": false
      }
    },
    {
      "type": "mcp",
      "server_label": "billing",
      "server_url": "https://mcp.example.com",
      "allowed_tools": ["searchInvoices", "createCreditNote"],
      "require_approval": "never"
    }
  ],
  "text": {
    "format": {
      "type": "json_schema",
      "name": "invoice_result",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "invoiceId": { "type": "string" },
          "total": { "type": "number" },
          "currency": { "type": "string", "enum": ["USD", "CNY", "EUR"] },
          "status": { "type": "string", "enum": ["paid", "open"] }
        },
        "required": ["invoiceId", "total", "currency", "status"],
        "additionalProperties": false
      }
    }
  },
  "service_tier": "fast",
  "safety_identifier": "billing-user-42"
}

Structured Outputs: shape, streaming, shared-origin schemas

Today’s Structured Outputs can pin a final object: on Responses it lives under text.format, same rules as Chat Completions response_format, different field names. The official pitch is type safety, detectable refusals, and less “please emit JSON” prompting. It does not sell meaning: locking shape is not locking business semantics. Truncated stream strings, silently dropped $ref, oversized schemas, and dynamic enums are still the usual 2026 integration bruises. Our Structured Output tutorial and AI JSON error guide said the same thing: the platform guarantees matching braces, not a correct invoice total.

The developer pain a keynote loves is this trio: (1) streaming Structured Output whose deltas are legal partial objects or JSON Patch, not broken strings; (2) an official way to declare tool schema and output schema as the same origin, so you stop hand-maintaining twins; (3) a less lossy JSON Schema subset—fewer silent drops of oneOf / $ref. Odds sit below a Completions sunset, but if it ships, your stream parser and schema directory get named on day one.

What you can do is stage-proof: keep schemas/output/ apart from schemas/tools/; write each output schema as Draft 2020-12 with strict + additionalProperties: false and a full required list; feed the same file to the platform and the local validator. When platform errors and local errors disagree, Diff the two schemas before you blame the model. The object you hand to a user or a downstream system must not share a file with a checkpoint or with MCP arguments—see AI Agent State.

Tool Calling: programmatic hops and multi-agent

Tool calling in 2026 is no longer “the model picks one function, you run it once, you stuff a string back.” On 9 July GPT-5.6 added programmatic tool calling, explicit prompt-cache controls, persisted reasoning, and multi-agent orchestration (Beta) on Responses. The model can chain tools structurally and treat child agents as first-class objects. Beta’s usual fate is a DevDay GA with limits and an SLA.

If it graduates, your orchestrator must redraw the boundary: which hops the platform owns versus which stay in your loop. A new JSON family appears—intermediate child-agent artifacts. Do not dump them into the same file as final Structured Output or MCP arguments. Stamp correlationId / runId / callId on every tool log now so GA-day reconciliation is possible. Function-tool arguments are still often a JSON string: parse and schema-validate before execution—the same class of problem as MCP tools/call.

Tool search, Skills, and hosted shell already live only on Responses. If DevDay makes “search then call” the default, you will regret stuffing eighty functions into tools[]. Split tool packs by domain; reuse an allowed_tools-style allowlist on function tools too. Keep parameter schemas small, enums short, and additionalProperties off. Details in MCP and JSON Schema.

MCP: remote tools, approvals, and connectors

From GPT-5.5, Responses can hang a remote MCP server: the model first emits mcp_list_tools, then picks a call. The console has OpenAI-maintained connectors; the 19 May Secure MCP Tunnel lets ChatGPT, Codex, Responses, and AgentKit reach on-prem servers via a customer tunnel-client—mostly as an enterprise motion. Remember the default: the platform asks approval before data leaves for the remote, and you see mcp_approval_request in output. Once you trust a server, set require_approval per tool or to never. Use allowed_tools when the catalog is large—listing alone burns context.

The gap is obvious: typical projects still cannot one-click a private MCP, and the connector catalog does not cover homegrown tools. The forecast marked “self-serve hosted MCP / connectors, tunnel as a project toggle” as high-odds. This guide asks for one odds-independent habit: MCP Tool.inputSchema must share origin with model tools[].parameters. Wrapper field names may differ; properties, required, and enums must be generated from one source file. The platform will not runtime-validate your server. The chain stays parse → schema → business rules. Remote MCP is stateless JSON-RPC; business position does not live on the server—see Stateless MCP.

Do not mix this with A2A. MCP is the model calling a tool; A2A is sideways agent delegation with its own task lifecycle and artifacts. OpenAI already consumes MCP at the model layer; sideways delegation is still a hole. If DevDay nods at AAIF / A2A, that is interop—not a reason to merge an Agent Card into inputSchema. See A2A vs MCP. When private-tool arguments drift from MCP params.arguments, Diff the generator, not the model’s mood.

Six JSON Schemas to freeze now

Do not ship one giant JSON. Split by consumer: the model reads the output shape, the runtime reads tool arguments, the MCP server reads inputSchema, the orchestrator reads session items and child-agent artifacts, the ledger reads usage. Six files cover the surfaces DevDay is most likely to touch. Author Draft 2020-12 sources; when you generate an OpenAI wrapper (text.format / parameters), add packaging only—do not edit properties.

File What it locks Who reads it Does an extra keynote parameter kill it?
schemas/output/invoice_result.jsonFinal Structured Output objectResponses text.format / local validatorNo. strict + additionalProperties:false still holds
schemas/tools/searchInvoices.jsonFunction-tool parametersResponses tools[] / execute-and-parseNo. GA orchestration does not reshape inputs
schemas/mcp/searchInvoices.jsonMCP Tool.inputSchema (same origin as the row above)MCP server and tools/callNo. Self-serve hosted MCP only changes provisioning
schemas/session/conversation-item.jsonConversations / output item unionExport, replay, complianceFields may grow; freeze the type enum first
schemas/agent/child-artifact.jsonChild-agent intermediate envelopeMulti-agent orchestratorNo. Beta→GA needs a separate file even more
schemas/obs/usage-record.jsonusage + cache + safety + request_idLedger and reconciliationNo. A new dashboard must map onto keys you already store

Keep an index in the repo that states which two files share origin and which file is only a wrapper. The index is ordinary JSON, easy to review and Diff:

{
  "schemaVersion": "1.0",
  "pack": "devday-2026-prep",
  "files": [
    {
      "id": "so.invoice_result",
      "path": "schemas/output/invoice_result.json"
    },
    {
      "id": "fn.searchInvoices.parameters",
      "path": "schemas/tools/searchInvoices.json"
    },
    {
      "id": "mcp.searchInvoices.inputSchema",
      "path": "schemas/mcp/searchInvoices.json",
      "sameOriginAs": "fn.searchInvoices.parameters"
    },
    {
      "id": "conv.item",
      "path": "schemas/session/conversation-item.json"
    },
    {
      "id": "agent.artifact",
      "path": "schemas/agent/child-artifact.json"
    },
    {
      "id": "obs.usage",
      "path": "schemas/obs/usage-record.json"
    }
  ]
}

The loop is always parse → schema → business rules, whether or not the model “gets smarter” on stage. Keep one real failure fixture each for truncated JSON, a schema error, and drifted MCP arguments. On announcement day you Diff new behavior against those fixtures—not against a social recap.

You can run it in the browser:JSON formatfor parse;JSON Schema validatorfor output and tools;JSON Difffor model arguments vs MCP params.arguments. Nothing leaves the machine.

Further reading: DevDay ten predictions, Assistants → Responses migration, Structured Output, MCP and JSON Schema, Stateless MCP, A2A vs MCP.

FAQ

Is this the official agenda?

No. OpenAI has published the date, venue, and a technical focus on APIs and developer tools. “Possible changes” on the four tracks are engineering judgment from the 2026 changelog. The room may ship a subset—or spend the hour on models and Codex. The schema list does not depend on the agenda.

How is this different from the 2 September forecast?

The forecast covers ten directions (including tiers, identity, observability, multimodal). This piece expands only the four lines that change JSON shape and names six files to freeze. Read them together: the forecast watches the stage; the guide changes the repo.

Still on Chat Completions—too late?

No, and migrating now is cheaper than the week after a sunset date. Move tool calling and Structured Output first, then attach Conversations. Do not wait for the keynote to build the wrapper. The Assistants article has the steps—Completions wrappers can follow the same item model.

If the predictions miss, are the schemas wasted?

No. Splitting output, tools, MCP, session, artifact, and usage files is hygiene Responses, MCP, and Realtime already need. An extra keynote parameter does not make additionalProperties:false wrong.

Summary and next steps

What matters at DevDay 2026 is not another model name. It is whether the platform collapses Responses, Structured Outputs, Tool Calling, and MCP into a default stack. All four tracks say the same thing: fewer parallel APIs, more JSON you actually have to honor.

Before 29 September, stop new Completions code, split the six schemas, keep a failure fixture. On the day, read the changelog—not a recap thread. When you need to check a payload shape, open JSONVue.