Analysis
Will Siri AI become an AI agent? Apple Intelligence, JSON, APIs, tool calling, and App Actions
Siri is learning to pick a tool, fill arguments, and read the result. That looks like an agent. It will not become an open loop you wire yourself — the catalog is the system’s, the contract is App Intents, the shape is typed parameters, and JSON appears when the action hits your server.
In 2026 “will Siri become an AI agent” is usually asked in the same breath as “will Siri become ChatGPT.” Both miss. Apple’s developer page now names the product surface Apple Intelligence: a personal system powered by next-generation Apple Foundation Models, stressing personal context, app actions, and on-screen awareness, and wiring your App’s content and actions into Siri AI. In engineering terms an agent is a model that picks tools in a loop, a runtime that executes them, and structured data that carries state. Siri in 2026 has the last two in outline — it can read your App Entities, act through Intents, and resolve “this” on screen. What it does not have is the open loop you already know: you cannot hang an arbitrary OpenAI tools[] array on system Siri, and you cannot let it patch files in your repo. This piece uses our AI agent definition, then separates Apple Intelligence, App Intents / App Actions, tool calling, and JSON, and ties them to Apple AI agents and APIs, the three-way comparison, and the Gemini strategy.
The short answer: it will act, not become an open agent
Use the engineering definition, not keynote adjectives. An agent needs a goal, perception, and action: the model chooses which tool and which arguments; the runtime executes and writes the observation back. Classic Siri mostly heard a phrase and opened a shortcut or a system capability, with almost no “look at the result, then decide” loop. Siri AI in 2026 starts to close that loop: it can pick an action from the App Intents catalog, resolve “save this to Notes” with on-screen context, and hand an entity to the next Intent across Apps. Behaviorally, it is becoming a system-level agent.
It is still not the open agent you stand up in an IDE or a support desk. The difference is who draws the tool boundary, not how clever the model is. An open agent’s catalog is yours: OpenAI tools[], Anthropic tool_use, MCP tools/list; arguments are almost always JSON. System Siri’s catalog is the Intents already declared on the device, preferably fitted to an App Schema; parameters are Swift types, not a blob you parse. The privacy fence differs too: on-device and Private Cloud Compute decide which context may leave the machine — not a line in your system prompt that says “you may read Contacts.” Treat Siri as “another ChatGPT you can hang functions on” and you ship the wrong migration: deleting App Intents, or calling the Gemini API from iOS “to match Cupertino.”
| Capability | 2026 Siri AI | Open coding / ops agent |
|---|---|---|
| Pick a tool | The system picks from App Intents / App Schemas | The tools[] or MCP list you inject |
| Fill arguments | Typed Intent parameters, checked at compile time | JSON arguments, parsed at runtime |
| Decide after the result | The system picks the next hop or stops; cross-App often needs on-screen context | You write result JSON back into the messages |
| Stop conditions | User confirm, system policy, privacy and permissions | maxSteps, Schema failure, policies you wrote |
One sentence: Siri AI will become an agent inside the system, not an agent runtime you own. What you control is the shape of the actions you expose, and the JSON after those actions hit HTTP. The next two sections split product from contract so “it can get things done” is not read as “an open tool loop.”
The three things Siri AI actually gained in 2026
Apple’s developer docs and the WWDC26 line fold Siri’s App Intents work into three capabilities. They explain why it feels like an agent — and why you still write Intents instead of pasting a JSON Schema and calling it done.
- Reach your content: model business objects as App Entities, fit them to an Entity Schema, and contribute them to the Spotlight semantic index. Only then can Siri find “that paid order” in personal context instead of launching your home screen.
- Take action: once an Intent fits an Intent Schema (commerce, photos, communication, and other pre-trained domains), natural language can hit perform() without a phrase list per wording. That is the app actions the official pages keep repeating.
- Read the screen: annotate views as entities with View Annotations or NSUserActivity so people can say “this” and “that.” Cross-App requests usually start here, not in a planner you wrote.
Stack those three and users feel “Siri can get things done.” Developers feel their App has become a toolbox for the system agent. WWDC26’s Build intelligent Siri experiences with App Schemas also freezes the test order: isolate Intent logic, then Shortcuts for parameter shape, then Spotlight for the index, and only then Siri end to end. Skip the first three and debug with voice only, and the bill gets large.
The model layer is moving too; do not read that as a product swap. The January 2026 joint statement hung the next AFM foundation on Gemini technology; in June, Private Cloud Compute expanded onto Google Cloud for heavier agentic tool-use. Users still see Siri / Apple Intelligence — not a Gemini badge, and not Google Assistant. For the strategy split see why Apple leans on Gemini; for how that sits next to the ChatGPT extension, see the three-way comparison. For you, what changes is how boldly the cloud model will chain actions — not the surname of your parameter table.
Three layers: Apple Intelligence, models, developer surfaces
Flatten “Siri became an agent” into one layer and every later decision tilts. Split at least three: the product people see, where inference runs, and which contract your code implements.
| Layer | What you see | What you should assume |
|---|---|---|
| Product | Apple Intelligence / Siri AI; Settings still shows Apple’s brand | Experience and confirms stay in the system, not in your chat UI |
| Model | On-device AFM plus PCC; the heaviest work can use the expanded cloud tier | Model trademark ≠ request protocol; fallback arguments must still validate |
| Developer surface | System → App via App Intents; in-app models via Foundation Models Tool | Two contracts in parallel — do not force one JSON field set onto both |
Foundation Models is a second line. It lets an App run a LanguageModelSession on-device (and on the PCC path the docs describe), call tools, and constrain structured generation. That is “your App is the agent runtime,” not “system Siri is the agent and hits your Intent.” Many Apps will ship both: App Intents to be found by Siri / Spotlight / Shortcuts; Tool conformances so an on-device model can check stock or fill a form. How those wires meet HTTP is in how an Apple AI agent reaches Apps and APIs.
Shortcuts glues the layers. Apple Intelligence can assemble multi-step automations from natural language; once you adopt App Intents, your actions join that ecosystem next to capabilities such as Use Model. A user can say “find paid orders, then write them into Notes” — the orchestrator is still the system, not a planner on your servers. Each Intent’s parameters must stand alone, because the system may run only one hop, or reorder them.
App Intents, App Schemas, and App Actions
App Intents are not “add voice to Siri.” They are the contract a third-party App signs with Apple Intelligence: you declare what content you can offer and what actions you can run; the system decides when to call. Rolled out with iOS 18 and, by 2026, feeding Siri, Spotlight, Widgets, Shortcuts, and the system agent at once. Teams that still treat App Intents as optional are misreading the front door of system AI.
An App Schema is the pre-trained shape of that contract. An Entity Schema tells the system this is an order, a photo, or a conversation, so it can enter the semantic index. An Intent Schema tells it this is find, open, or share, so natural language can hit the action without a phrase corpus you maintain. Xcode macros generate conforming stubs and check them at compile time. Miss the domain and Siri may still treat your action as an ordinary shortcut, not a tool the system agent prefers. WWDC26 also added longer-running Intents, cancellation, and entities that sync across devices — more loop-like, still scheduled by the system.
“App Actions” is the phrase that drifts. On Apple’s developer pages, app actions are capabilities exposed through App Intents and executed by Siri AI and other system surfaces. They are not Google Assistant’s historic App Actions protocol, and not Android intent filters. Both products “let an assistant call an App”; the fields, catalogs, and privacy models do not match. In design docs write “App Intents / App Schema” and leave “App Actions” for user-facing copy. Mix the names and a backend engineer will hunt a Google-style JSON-RPC that does not exist.
Tool calling: the system agent vs in-app Tools
Industry tool calling / function calling is one mechanism: the model does not touch the database; it emits “please call this function with these arguments,” and the runtime executes. In 2026 the shells differ; arguments are still a JSON object. On the Siri side the same idea becomes a typed Intent: the system model picks findOrders, fills email / status / limit, and enters your perform(). If perform() hits your HTTP API, you then encode those fields as JSON. Do not share one “just parse it” path across both hops.
Store golden payloads at the intent layer, not under a model trademark. The JSON below only says who started, which schema, which action, and which arguments. runtime.onDevice is fine in logs. Do not put it in business validation — on-device fallback and the full cloud tier must accept the same arguments.
{
"source": "siri-ai",
"schema": "commerce.findOrders",
"intent": "findOrders",
"arguments": {
"email": "ada@example.com",
"status": "paid",
"limit": 5
},
"runtime": {
"surface": "siri",
"onDevice": false
}
}
The same business fields, in an open agent you own, wear a vendor tool_calls shell. arguments is often a string: JSON.parse first, then validate against the same Schema. When Siri calls in you never see that shell — only typed parameters. You see JSON again when you leave the device for your own API.
{
"id": "call_8f3a",
"type": "function",
"function": {
"name": "findOrders",
"arguments": "{\"email\":\"ada@example.com\",\"status\":\"paid\",\"limit\":5}"
}
}
The contrast is sharp: the name is findOrders either way; the keys should share a Schema. The system path has no tool_call id; the open path has no App Schema domain. Do not treat Gemini generateContent, OpenAI tools[], and an App Intent parameter table as default values for each other. Structured Output governs the final answer, not this hop’s inputs; why the files split is in AI Structured Output.
JSON, APIs, and local validation
As Siri looks more like an agent, failures do too: extra hops, optional fields filled in, wild enums, two keys missing on an on-device fallback. A stronger model does not retire validation; it makes misses cost more. Every egress hop is still parse → Schema → business rules.
- Encode Intent parameters as JSON and parse immediately; on failure log the intent name and raw fields and return a retryable error — do not dump a Swift exception at the user.
- Pin types, enums, and required with one Draft 2020-12 Schema. Prefer additionalProperties: false so extra keys from the cloud tier do not leak downstream.
- Business gate: permissions, foreign keys, date ranges. Only then hit the order service. When the system agent retries, your API must be idempotent.
Line up three blobs: the parameter object Siri / Shortcuts sent, the HTTP body you emit, the object the server actually used. Mismatch is almost always the adapter. In the browser use JSON formatter to see fields; JSON Schema validator to pin types; JSON Diff to see which keys the on-device fallback drops versus the cloud tier. Share valid / missing-field / wrong-enum fixtures in CI and on-device Siri debug.
Further reading: Apple AI agents and JSON, what an AI agent is, AI JSON errors guide, MCP and JSON Schema.
FAQ
Is Siri already an AI agent?
By the engineering definition it already has the outline of a system-level agent: it picks App actions, fills parameters, and uses on-screen context to decide again. It is not an open agent runtime — the catalog, stop conditions, and privacy fence belong to the system, not to a tools[] list you inject.
Are the App Actions in this article Google Assistant’s App Actions?
No. On Apple’s pages, app actions are capabilities App Intents expose to Siri AI. Google’s historic App Actions are a different assistant integration. The contracts, fields, and catalogs are not interchangeable. Internally write App Intents / App Schema.
Should we call OpenAI or Gemini tool calling from iOS to “match” Siri?
When the system calls in, keep App Intents. For an in-app on-device model, use Foundation Models Tool. Use those vendors’ APIs only when your own servers need their models. Do not mix the three JSON field sets.
Can Siri still call my App if we skip App Intents?
It can launch the App. It cannot reliably get work done. Without Entity / Intent Schemas, Siri lacks indexable content and executable actions, and “this” across Apps has nothing to bind to. Treating App Intents as a voice extra in 2026 is opting out of the system agent’s toolbox.
Summary and next steps
Siri AI will become an AI agent — inside the system’s fence. Apple Intelligence supplies personal context, app actions, and on-screen awareness; the model layer can borrow Gemini technology and a heavier PCC tier; developers still hand over actions through App Intents and run in-app loops with Foundation Models Tool. The open tool-calling JSON shell will not appear on the system request. It will appear on your own API.
Next is concrete: list the App Schemas you will adopt, write Intent parameters and the HTTP body against one Schema, and replay valid / missing-field / wrong-enum in JSONVue. Model trademarks will move. Action shape should not.