Tutorial

How does an Apple AI agent connect apps, APIs, and tools? What role does JSON play?

The system agent uses App Intents; the in-app model uses Foundation Models tools. JSON is not decoration — it is the shared shape for argument contracts, structured output, and HTTP bodies.

“Apple AI agent” is often treated as one product name. In production you hit at least two different call chains: Apple Intelligence picking an app and an action for the user, versus your app running an on-device model that decides which tools to call. Both touch apps, APIs, and tools, but who owns the session, who generates arguments, and where JSON appears are not the same. Split those layers or schemas, logs, and debugging have nowhere to land.

First: who is running the model?

Start with who initiated this inference. If the person talks to Siri or system-level Apple Intelligence, the model runs on the system side (on-device or Private Cloud Compute). Your app is only the capability being routed to. You expose App Intents: typed actions, entities, and parameters. The system agent decides when to call; your code performs. Official entry: App Intents documentation.

If they are already inside your app and you start a Foundation Models LanguageModelSession, that is a different chain: your process drives the model and you inject the tools. The system will not pick “which app to open” because the session already lives in the app. Tools can read contacts or calendar, hit your own network API, then write results back into the transcript. Framework notes: Foundation Models. For session-level tool calling, see WWDC25 Meet the Foundation Models framework.

A third pattern is now common: an external host (Claude, ChatGPT, a home-grown agent) calls your service over MCP or plain HTTPS. The model is not on Apple’s stack, but the payload is almost always JSON. The same domain functions can serve App Intents, Foundation Models tools, and HTTP — only the adapters differ.

Call chain Who runs the model What the app provides
System agent Apple Intelligence App Intents / App Entity
In-app agent Your LanguageModelSession Tool protocol and @Generable arguments
External agent Third-party host JSON-RPC or REST JSON

Do not collapse the three chains into one “universal agent” diagram. When debugging, identify the chain first: system routing failures belong in Intent declarations and parameter types; wild in-app tool calls belong in name, description, and the generable Arguments shape; mismatched external APIs belong in HTTP JSON and schema. Mix them and it just feels like “the model is being random”.

App Intents: how the system agent reaches your app

To the system agent, an app is not “open it and look around” — it is a discoverable catalog of capabilities. You declare action names, natural-language descriptions, parameter types, and results with AppIntent. Spotlight, Shortcuts, Siri, and Apple Intelligence share that catalog. When someone says “mark this invoice paid”, the system must map the utterance to your MarkInvoicePaid, not let the model tap around the UI.

Parameters are the join. The system turns speech into typed values: enums, dates, AppEntity references. In Swift you see structs, not prose. Across processes or frameworks that structure still needs a serializable shape — debugging, logs, and server replay end up as JSON or an equivalent property list. Design Intent parameters as a field set that can be written as a JSON object and API adapters get much cheaper.

A common tradeoff: should the Intent hit the network? Short actions can finish inside perform. Once you have auth, pagination, or idempotency, perform should only validate parameters and call your domain service; the service sends the JSON request. The system agent does not need your REST path — it needs success or failure semantics. You need the path, because billing, audit, and retries live at the API layer.

Foundation Models: how the in-app model calls tools

The in-app agent connection looks more like handing the model a function spec. You implement Tool: give it name, description, plus an @Generable Arguments. The framework puts that into the prompt so the model can decide when to call. On a call it first generates arguments, then the framework runs call(arguments:) and inserts the return value (usually a String or a generable type) back into the transcript. The model writes the final answer from that. You are not asking it to invent URLs — you are asking it to choose among tools you allowed.

Argument generation is structured output, not “please answer in JSON”. @Generable and dynamic schemas pin fields, enums, and nested objects at decode time. In Swift you get typed instances; if you persist, log, or hand them to URLSession, encode to JSON then. Apple summarizes the capabilities as language understanding, structured output, and tool calling in the Foundation Models overview.

struct FindOrders: Tool {
  let name = "findOrders"
  let description = "Find recent orders by customer email and status."

  @Generable
  struct Arguments {
    var email: String
    var status: String
    var limit: Int
  }

  func call(arguments: Arguments) async throws -> String {
    // Call your domain API, then return a compact summary.
    return "3 orders, latest is paid"
  }
}

In that snippet the model does not invent field names. It must fill email, status, limit. The description decides whether it gets selected: too broad and it fires on everything; too narrow and it stays silent when you need it. Keep tool output small — dumping a full order JSON into the transcript burns the context window. Return a summary, then call a second tool for an id when you need detail.

Tools can chain. The model uses the first tool’s output as the second tool’s input; the framework runs them in order. Enforce idempotency in the domain layer: marking the same orderId paid twice must not charge twice. The agent loop cannot see your database constraints. Valid JSON arguments are not the same as a correct business outcome.

What role JSON plays

Swift types are the compile-time contract; JSON is the runtime contract. The moment the agent leaves the process — backend calls, files, another model, test fixtures — the shape must become language-agnostic text. JSON plays at least three roles on that chain. Mix two of them and you get “it parses, but every field is wrong”.

1. The interchange format for tool arguments

Foundation Models represents structured values internally as GeneratedContent. The most useful debug view is often “what this Arguments value looks like encoded as JSON”. Names, optionality, array vs object must stay stable. Log arguments as JSON so they line up with production API access logs: did the same customerId actually reach the server?

{
  "email": "ada@example.com",
  "status": "paid",
  "limit": 5
}

2. The schema for structured output

Even when you skip tools and only want an invoice summary, you still need a schema: required keys, enum values, number vs string for money. Apple’s side uses @Generable; cloud models usually use JSON Schema. When both describe the same domain object, share one field table so the app does not say totalCents while the API says amount. Shared rules for writing schemas: Understanding JSON Schema.

3. The payload from app to API

Once a tool’s call hits the network, JSON is the HTTP body. The agent does not replace API design: auth headers, idempotency keys, and error objects are still yours. The model only fills business fields; transport, pagination, and rate limits stay ordinary backend work. Shape errors as stable JSON (code, message, retryable) so the model can decide whether to switch tools or explain the failure.

All three layers can share one schema document: Intent parameters ⊂ tool Arguments ⊂ HTTP body. Subsets make tests easy: one valid JSON fixture, then API, then Tool, then the Intent adapter. A superset (HTTP has three extra internal fields) is fine, but keep those fields off the model or it will “helpfully” fill keys you never meant to publish.

Wiring HTTP APIs and MCP

When an in-app tool calls REST, encode explicitly. Do not send GeneratedContent as Data. Map to your Codable model, then JSONEncoder. You then own enum raw values, date formats, and key strategy (snake_case). The model produces domain values; the encoder owns protocol details.

{
  "tool": "markInvoicePaid",
  "arguments": {
    "invoiceId": "inv_9f2",
    "paidAt": "2026-08-20T09:00:00Z"
  }
}

MCP turns “tool name + argument object” into JSON-RPC. For Apple developers that is just a third adapter: the same markInvoicePaid(invoiceId:paidAt:), App Intent goes through perform, Foundation Models through Tool.call, MCP through tools/call. Do not rewrite business logic for MCP. External agents are more likely to send the wrong types (numbers as strings), so the server must still validate — “it is already JSON” is not trust.

Write privacy boundaries into tool descriptions. An on-device model that can read the calendar is not a license to POST events to your server. Tool output for the model can be a local summary; build upload JSON only when the user clearly wants sync. Review data-egress policy separately for the system agent and the in-app agent, and make the source visible in logs.

How to inspect JSON when you ship

When integrating an agent, another prompt is less useful than three snapshots side by side: the model’s argument JSON, the HTTP JSON you sent, and the JSON the server returned. If the shapes disagree, the bug is almost always the mapping layer, not “the model isn’t smart enough”. Format, validate, and diff locally — faster than staring at Optionals in the Xcode console.

Walk a sample through the browser first: JSON formatter to confirm it parses; JSON validator to catch trailing commas and bad types; pin shared Intent / Tool / API fields with JSON Schema; then use JSON Diff to diff “model output” vs “actual request body”. Cloud structured output differs in API, but “contract first, then parse” is the same — see How the Gemini API outputs JSON.

Keep three fixture files: tool-args.valid.json, http-body.valid.json, http-error.json. In CI, validate the first two against the same schema. Use the error fixture to check the agent tells the user an actionable failure instead of swallowing retryable: true.

FAQ

Can App Intents and Foundation Models tools share the same parameters?

Share the domain model, but do not assume the system can run your Tool protocol. Intents are for system discovery and permissions; tools are for the current session prompt. Put a mapping in the middle. JSON fixtures test the domain model, not one framework type.

Why not let the model emit a full HTTP request?

URLs, headers, and signatures must not be invented by the model. It fills business fields; the client sends a fixed template. One hallucination otherwise hits the wrong environment or the wrong credentials.

Do JSON and @Generable conflict?

No. @Generable is Swift-side generation and decoding; JSON is the cross-language, cross-network shape. Stabilize a field table first, then generate Swift macros and JSON Schema from it.

Must tool return values be JSON?

Not necessarily. The model can see a short text summary; the server must get JSON. Do not mash both into one string and regex-parse it.

Summary and next steps

Agents on Apple’s stack are not one socket. They are three composable wires: the system agent discovers and calls the app via App Intents; the in-app model calls your code via Tool; an external host calls the same domain service with JSON. JSON makes arguments, schemas, and API bodies speak one shape. Types cover compile time, schemas cover runtime, business checks cover truth.

Next: list three to five domain actions. For each, write a minimal JSON object and a schema, then decide whether it appears on an Intent, a Tool, or HTTP. Stabilize the shape before you add prompts and multi-tool orchestration.