Analysis

Why is Apple leaning harder on Gemini? What actually changed in Apple’s 2026 AI strategy?

Leaning on Gemini is not handing Siri to Google Assistant. What changed is where the foundation model comes from and where the heaviest inference runs. What did not: on-device first, the PCC privacy bar, and the JSON shape you expose.

On 12 January 2026 Apple and Google published a short joint statement: the next generation of Apple Foundation Models would be based on Google’s Gemini models and cloud technology, to power future Apple Intelligence—including a more personalized Siri later that year. The same note said Apple Intelligence would keep running on Apple devices and Private Cloud Compute, with Apple’s privacy standard intact. Six months later Apple’s security engineers dropped the other shoe—PCC left Apple’s own data centers for the first time, running the heaviest Apple Intelligence workloads on Google Cloud with NVIDIA. Stack those two facts and headlines collapse into “Apple defected to Gemini.” The sharper reading: Apple outsourced foundation-model capacity to the partner that could absorb scale, and kept inference control, the product surface, and the privacy story. If you ship Apps or APIs, the question is not whether Siri becomes Google Assistant—Craig Federighi shut that misread down after WWDC—but whether stronger cloud models will emit richer tool-call JSON than your Schema can survive. This piece separates the official record, the PCC expansion, and reported figures, then ties them to our Apple AI Agent, Gemini JSON output, and Structured Output guides.

The two things that actually shipped in 2026

Keep “on the official page” apart from “in a leak.” The 12 January joint statement is two paragraphs. The first is scope: a multi-year collaboration; next-gen Apple Foundation Models based on Gemini models and cloud technology; future Apple Intelligence, including a more personalized Siri. The second is rationale and fence: after careful evaluation Apple called Google’s AI the “most capable foundation” for AFM, and restated that Apple Intelligence stays on-device and on PCC. It does not name a dollar figure, a parameter count, a replacement for Writing Tools, or a user-facing Gemini brand.

On 8 June, Apple Security Research published Expanding Private Cloud Compute. The claims are hard: PCC reaches third-party data centers for the first time; Apple works with Google and NVIDIA to run new Apple Intelligence workloads on NVIDIA GPUs in Google Cloud; the model family spans on-device to cloud; the most demanding jobs—agentic tool-use and complex reasoning—go to PCC on Google Cloud. The five PCC requirements did not move: stateless computation, enforceable guarantees, no privileged runtime access, non-targetability, verifiable transparency. What is new is the implementation: NVIDIA Confidential Computing, Intel CPUs with TDX, Google’s Titan chip, plus a cryptographically verifiable, append-only ledger of PCC hardware on Google Cloud. Apple is explicit: wherever the iron lives, PCC software stays under Apple’s control; devices trust only PCC software Apple has cryptographically approved.

Thread the timeline. In 2024 PCC ran only on Apple silicon, covering inference the on-device model could not. In 2025 Apple Intelligence launched with roughly hundred-billion-scale cloud models that still felt in-house, while a deeply personal Siri slipped. In January 2026 the “foundation” of the next AFM moved onto Gemini. In June 2026 the heaviest cloud tier landed on Google Cloud GPUs—still wrapped in PCC. The strategy shift is not a sudden fondness for Google. It is an admission: personalized Siri plus agent-grade tool calling, for two billion active devices, was more than a homegrown cloud model and Apple-only halls could deliver on the 2026 clock.

Why the Gemini dependence deepened

Apple is not model-less. The Foundation Models framework already lets an App run a LanguageModelSession on-device, constrain structured output with @Generable, and call tools—WWDC 2025 put that on the table. What it lacked was a frontier cloud model that could carry system Siri, and a hall that could carry peak inference. Embedding ChatGPT in Siri in 2025 was a user-visible side door: hard questions could go to OpenAI, but that was not the AFM base. The 2026 Gemini deal sits on a different layer. It enters AFM training and foundation. It is not another “Ask ChatGPT” pane.

Why Google rather than doubling down on OpenAI? In public Apple only says “most capable foundation.” Three product-boundary readings fit better than fan theories. First, Google will disappear into infrastructure: users see Siri and Apple Intelligence, not a Gemini badge, and not Google Assistant. Federighi later said the amount of Google Assistant they use is none. Second, Google can sell both model tech and cloud tech—the statement puts “models and cloud technology” in one sentence; the June PCC expansion cashed that check. Third, the deal is multi-year and widely read as non-exclusive: Apple can still build inference silicon and buy capability. Reported figures—about a billion dollars a year, a custom ~1.2T-parameter Gemini—are unconfirmed by either side. Treat them as scale color, not an architecture input.

A motive headlines skip: time. Personalized Siri was supposed to arrive earlier and hit a dual gap—model quality and server performance. Later reporting said engineers tried Gemini-class loads on existing M-series servers and came up short, which is why the top tier moved onto NVIDIA GPUs in Google Cloud. Apple is also pushing a server chip, Baltra, with Broadcom, aimed at inference rather than training; outside timelines for volume have slipped. Read it simply: Gemini is a bridge, not the destination. The far bank is Apple’s own inference silicon and halls. Until that bridge is crossed, dependence will look deeper—because the hardest queries all go Cloud Pro.

How to read the three inference tiers

Do not flatten “Apple uses Gemini” into one runtime. By summer 2026 the public map is roughly: light work stays on-device with Apple Intelligence, no network required; mid cloud work stays on PCC on Apple silicon—AFM 3 Cloud, ADM 3 Cloud for images; the heaviest agent tool-use and hard reasoning go to AFM 3 Cloud Pro on expanded PCC on Google Cloud. All three are Apple Intelligence to the user. Any of them can hit your App Intents or your HTTP API. What changed is who invents the arguments in the cloud. What did not is the shape those arguments should have.

Tier What typically runs What you should assume
On-deviceApple Intelligence on-device; short commands, visible contextFewer fields, more conservative; degrades offline
PCC · Apple siliconAFM 3 Cloud, ADM 3 Cloud, mid-tier cloud jobsStill Apple halls; JSON shape matches the system agent
PCC · Google CloudAFM 3 Cloud Pro: agent tool-use, hard reasoningMore willing to call tools and fill optional fields; Schema must hold

The table is a debugging lens. If a user says Siri both edited a calendar and checked inventory, Cloud Pro likely lit up: longer tool chains, fatter argument objects. If they say it dies offline, that is on-device fallback, not “Gemini is down.” You do not need to know which GPU served the request. You do need to assume the cloud model will call more tools and fill more optional keys, so extra keys, type drift, and wild enums show up more than in 2025. Local Schema checks are not leftover ceremony. They are the layer the new stack amplifies.

Dependence is not giving up the product

“Apple sold its soul” is a headline, not a reading of the documents. The joint statement locks the experience to Apple Intelligence. The PCC essay locks control to Apple-signed software. The post-WWDC line locks Google Assistant usage at zero. Google gets distribution, license revenue, and a proof that Gemini can be someone else’s base. Apple gets time: a trained frontier model to cover the Siri it could not ship in 2025–2026, without dropping the privacy pitch.

Three things actually changed hands: where the next AFM’s teacher signal / foundation weights come from; where the heaviest inference hall sits; how much peak compute is rented on NVIDIA GPUs. Four things did not: how the system agent picks an App (still App Intents); how an in-App model calls tools (still Foundation Models Tool); how your backend accepts HTTP JSON; the brand in Settings. Mix those layers and you ship the wrong migration—calling Gemini API from iOS “to match Cupertino,” or deleting App Intents “because it’s all Google now.” Both mistake an infrastructure deal for a product swap.

The privacy story was not officially retired; the implementation got more winding. PCC moved from “trust only our silicon” to “trust a verifiable third-party confidential-computing pipeline.” Security researchers will probe that. Apple’s answer is published binaries, research-mode nodes, the bounty program, and a summer preview that ramps toward the full protection set. For PMs: you can still say user data is not kept on servers by default, if you can explain stateless + attestation + short TTL. For engineers: logs should not hold replayable user prose, and tool results should be cropped to the minimum fields.

For developers: the JSON contract did not change owners

Return to the contract you control. When system Apple Intelligence hits your App, it still arrives through App Intents: typed parameters, not a blob of Gemini prose you parse. In-App on-device models still use Foundation Models: @Generable Arguments, then you decide in call(arguments:) whether to JSON-encode an HTTP body. Both chains are alive in 2026. Gemini changes how boldly the system cloud model will chain tools. It does not rewrite your parameter table into another protocol.

The other mix-up: your own backend calling the Gemini API. That is Google’s developer product, with response_format / JSON Schema, and it is not the Apple Intelligence request path. We already covered that chain in Gemini API JSON output. “Apple used Gemini tech to train AFM” does not mean “Siri on the iPhone now speaks generateContent JSON.” From the device, egress is a PCC envelope. On your servers you only see Google field names if you chose to call Gemini.

In practice, store golden payloads at the intent layer, not under a model trademark. The JSON below only says who started, which tool, and which arguments. runtime.pccTier is fine in logs. Do not put it in business validation—Cloud Pro and on-device must accept the same arguments Schema, or one fallback will break order lookup.

{
  "source": "apple-intelligence",
  "intent": "findOrders",
  "arguments": {
    "email": "ada@example.com",
    "status": "paid",
    "limit": 5
  },
  "runtime": {
    "onDevice": false,
    "pccTier": "cloud-pro"
  }
}

Once you have live traffic, drop pass and fail samples into JSONVue: format to see fields, JSON Schema to pin types and enums, Diff to see which keys the on-device fallback drops versus Cloud Pro. Apple leaning harder on Gemini in 2026 does not retire that step. Stronger cloud models fill more fields and call more tools; local validation is worth more. Further reading: how an Apple AI Agent reaches Apps and APIs, and AI Structured Output.

FAQ

Is Siri Google Assistant now?

No. The joint statement and the post-WWDC line keep the experience on Apple Intelligence / Siri. Gemini tech enters the next AFM foundation and the heaviest cloud hall. It does not replace the system assistant with Google’s product name.

Should my App switch to the Gemini API?

If the system agent is calling you, stay on App Intents. If your App runs an on-device model, stay on Foundation Models. Use the Gemini API only when your own servers want Google’s models. Do not mix the two JSON field sets.

Can I trust the $1B / 1.2T-parameter reports?

Treat them as unconfirmed scale reporting. Neither number appears in the joint statement or the PCC essay. Architecture should rest on three public facts: multi-year collaboration, PCC on Google Cloud, Cloud Pro for the hardest jobs.

Is the privacy promise void?

Apple did not withdraw the five PCC requirements. It extended the implementation onto third-party confidential computing. Whether that is enough is a question for published binaries, research nodes, and the summer preview. In engineering, still design tool results as stateless and minimal.

Summary and next steps

Apple’s 2026 AI shift compresses to one line: borrow Gemini at the model layer, rent Google Cloud GPUs at the inference layer, keep Apple on the product and attestation layers. Dependence deepened because personalized Siri and agent tool-calling exposed the gap in the homegrown cloud stack. That is not handing over the assistant brand.

Next step is concrete: make App Intents parameters and HTTP tool bodies share one Schema, validate pass/fail samples in JSONVue, and assume Cloud Pro will call one extra tool and fill two extra optional fields. Model trademarks will move. The JSON shape should not.