Tutorial
What is a 1M token context window? GPT vs Claude vs Gemini in 2026
“One million tokens” is not a synonym for a smarter chat. It is the budget a model can see in one request. A bigger budget can hold a repo, a handbook, or a long recording. Max output, price tiers, and real recall do not grow automatically with the sticker.
In 2026 almost every flagship model card says 1M tokens. Marketing copy reads “ingest the whole book” or “drop in the entire repo.” Review meetings get stuck on a different set of numbers: what this call costs, whether the answer will truncate, and whether the model still remembers the middle 400K. This article treats 1M tokens as an engineering limit — the capacity of the context window, not a certificate of understanding. When you finish you can estimate how many tokens your materials use, compare GPT, Claude, and Gemini on window, output, and surcharges, and decide which JSON belongs in the window and which should be checked in the browser first.
What 1M tokens means: get the unit right
A token is the smallest unit the model uses to read, write, and bill text. It is not a character and not a word. A common English word is often a bit more than one token; Chinese, code, JSON keys, Base64, and indentation all convert differently. The same OpenAPI file can count differently across vendors and tokenizer versions. So “1M tokens” is first how many slots this model allows on this request — not “a million Chinese characters” or “a million English words.”
| Window size | Rough intuition (order of magnitude, not a promise) |
|---|---|
| 1 token | About 0.7–1 English word; often 1–2 Chinese characters; code and JSON fragment more |
| 8K–32K | A long article, an API note, a short conversation |
| 128K–200K | A large share of a book, or a mid-size repo plus a few agent turns |
| 1M (about 1.00–1.05 million) | A thick handbook + a multi-file repo + long history; newer Claude tokenizers often fit less raw text in the same 1M |
The context window is a shared budget for input + output + reasoning/thinking tokens (when they are metered). System prompts, tool schemas, retrieved snippets, chat history, and the answer being generated all draw from the same well. A 1M window does not mean you can load a 1M-token repo and then emit a 128K report — output counts against the window, and each vendor also sets a separate max output.
Treat 1M like disk capacity: the slots exist; random access to a page in the middle is not equally reliable. Finding a needle in a haystack is not the same as multi-hop reasoning, cross-file comparison, or noticing a missing field in a long JSON. Ask first “what must be visible at the same time,” then “what stable shape must come out.”
What a long context window is actually for
Jumping from 128K to 1M changes whether you still have to chunk first. Material that used to need retrieval or a summary can now go in as one pack. The jobs that actually pay off usually look like this:
- A whole repo or most of one: follow a call chain across files, compare two implementations, see the interface and the fixtures in one request.
- Long documents and contracts: policy manuals, audit workpapers, multi-attachment PDFs. The question needs both ends of the text present, not a single keyword hit.
- Large multimodal packs: Gemini-class APIs can count text, images, audio, video, and PDFs in one window. Hours of audio or a long video no longer have to be sliced first.
- Long agent traces: tool defs, many tool results, and failed retries stay in the window so the next hop can see the full observation. For the loop itself see
- Compare-and-extract: put two OpenAPI versions, two export JSONs, and one schema in the same request for a diff-grade reading, then emit structured deltas via Structured Output.
These jobs share one property: the answer depends on seeing several artifacts at once. If the task is “find one sentence in a hundred thousand pages,” retrieval is usually cheaper and steadier. A 1M window reduces bad chunk boundaries; it does not retire information architecture.
One overrated pattern: paste full site logs, a raw database dump, or unsanitized production JSON into the prompt. The window is large; so are the leak surface and the bill. Only material you have already decided may leave the machine should enter the window. Strip secrets locally, fix malformed JSON, then talk about “feeding it in one shot.”
Sticker window ≠ effective window: input, output, price
The 1M on the product page is an acceptance ceiling: go over and the request is rejected or truncated. It is not “quality at 1M equals quality at 8K.” On RULER-style evals, effective length is often about half the sticker; in production, cross-file work often slips after 200K–500K and mid-prompt instructions get dropped. Treat the advertised window as a hard roof. Treat your own regression set as the effective window.
The second ignored number is max output. GPT-5.6 Sol / Terra and Claude flagships allow about 128K on a sync request. Gemini 3.1 Pro documents 65,536, and the default maxOutputTokens is often lower — forget to raise it and you silently truncate. “Read the handbook, then write a same-length spec” may fit the window and still blow the output cap. The budget below treats output and headroom as first-class, not leftovers:
{
"window": 1050000,
"budget": {
"system": 2400,
"tools": 1800,
"repoPack": 420000,
"docs": 180000,
"history": 80000,
"reservedOutput": 128000,
"headroom": 238800
},
"rule": "never fill to the sticker; reserve output plus 20 percent headroom"
}
The third number is money. OpenAI states that for GPT-5.6 Sol / Terra, once input exceeds about 272K tokens, the whole request is billed at 2× input and 1.5× output. Gemini long context often has another surcharge band (docs commonly draw the line above 128K or 200K — check the live price list). Claude Sonnet 5 docs present 1M as the default window at standard rates, with no separate “enable long context” switch — but a finer tokenizer can turn the same text into more tokens, so effective cost still rises. Prompt cache helps repeated prefixes; it does not help a pack that changes every call.
2026 comparison: GPT, Claude, Gemini
The table below follows public September 2026 model cards for flagship tiers. It does not cram every snapshot into one cell. Numbers move; the live docs win. The axes that actually fork a choice are window, output, surcharge, and modality.
| Axis | GPT-5.6 Sol / Terra | Claude Sonnet 5 / Opus 5 | Gemini 3.1 Pro |
|---|---|---|---|
| Context window | 1,050,000 | 1,000,000 (Haiku 4.5 stays at 200K) | 1,048,576 |
| Max output | 128,000 | 128,000 (higher Batch tier exists) | 65,536 (defaults are often lower) |
| Long-context billing | Input >272K: whole request 2× input / 1.5× output | Docs: 1M is default, no separate surcharge band; newer tokenizer is finer | Common surcharge above 128K or 200K — confirm the official list |
| Input modalities | Text, image | Text, image (PDF via platform features) | Text, image, audio, video, PDF |
| Same-family caveats | Luna is about 400K — do not treat the alias as the same capacity | Same 1M holds less raw text after the tokenizer change; recount | Widest multimodal; shorter output ceiling — budget generation first |
| Best fit | Long reports, Responses toolchains, jobs that need 128K out | Long-document reasoning, agent loops, read-then-answer stability | Repo + audio/video/PDF in one shot; short or schema-shaped answers |
Read the model cards, not only third-party tables. OpenAI’s GPT-5.6 Sol lists the 1,050,000 window, 128K output, and the 272K surcharge line. Anthropic’s Claude Sonnet 5 overview treats 1M / 128K as default. Google’s Gemini 3.1 Pro Preview lists 1,048,576 input and 65,536 output. Prices and snapshot names change; integration code should read usage fields instead of last year’s constants.
Pick from the bottleneck. If you must emit nearly 100K tokens of a migration note or a long JSON array, the output cap matters more than the window — GPT and Claude flagships are wider. If one request must hear a meeting recording and also see a design deck and a repo, Gemini’s modalities are a real difference, not an adjective. If you will do multi-hop edits inside 1M and cannot drop mid-prompt instructions, do not trust the sticker — measure effective recall with your own “three needles, two contradictions” set, not a single haystack demo.
All three support tool calling and Structured Output; the wrapper fields differ, and you still validate JSON locally. For Gemini’s split between “just JSON” and “fields from a schema,” see the Gemini API JSON output guide — this article will not repeat the SDK.
When to fill the window, when to retrieve
A 1M window does not retire RAG. It pushes the chunk boundary back. Four rules are enough:
- The answer must see A and B together (two specs, caller and callee, schema and instance) → prefer one window. Do not summarize each side separately.
- The answer is “locate a small span in a huge corpus” → retrieve first, then send hits plus metadata. Do not pay a 272K surcharge to skip a search pipeline.
- The pack changes every request (new diff, new export) → long context is expensive. Split cacheable prefixes (system prompt, tool list, stable manual) from the volatile pack.
- The material must not leave the browser (secrets, production user fields) → a larger window is not a reason to upload. Clean, validate, and sample locally, then decide whether a cloud model sees it.
Agent loops fill a window fast: every tool result stays. Summarize old observations before maxSteps; do not treat 1M as an infinite log. How the loop works: What is an AI agent.
Even when you choose “the whole pack,” pack it. Do not spend 1M on node_modules and minified bundles. An explicit context pack is auditable; “drag the repo in” is not:
{
"pack": "context",
"files": [
{ "path": "openapi.json", "tokens": 12000, "role": "contract" },
{ "path": "schema.ticket.json", "tokens": 400, "role": "outputShape" },
{ "path": "fixture.valid.json", "tokens": 800, "role": "example" }
],
"omit": ["node_modules", "*.lock", "generated/**", "minified bundles"]
}
The outputShape in the pack should be the schema you will actually parse — not whatever keys the model invents. How to pin shape: AI Structured Output.
JSON in long context: validate locally first
A long window increases the blast radius of bad input. A truncated export, a tool receipt with exploding extra fields, two schemas that share a name and not a meaning — after 400K tokens you will not find them by eye. The model will keep writing from the broken JSON, and you still pay. Fix three steps before upload:
- Parse: every file that will enter the window is real JSON / JSON5, not a log that “looks like” JSON or a cut-off export.
- Schema: validate against the contract you want the model to keep. With several files, confirm they still share one source of truth.
- Diff: compare old/new fixtures, model arguments vs downstream body, or two OpenAPI versions locally before you ask the model “why” or “how to change it.”
You can do this in the browser without uploading first: JSON formatter for parse and structure; JSON Schema validator for fields and enums; JSON Diff for two artifacts that will share a window. Files that fail validation do not enter the pack.
Long JSON from the model can still be syntactically broken or shape-drifted — see the AI JSON errors guide. Tool arguments and the final answer are two contracts. Sharing a window budget is not the same as having validated them.
FAQ
Is 1M tokens a million Chinese characters?
No. Tokens are tokenizer fragments. Chinese, English, code, JSON, images, and audio all convert differently; a vendor tokenizer change can also change how much text fits in the same 1M. Count your materials with that model’s tokenizer API. Do not substitute “pages × a rule of thumb.”
If every window is 1M, does the vendor still matter?
Yes. Max output, long-context surcharges, modalities, and effective recall diverge. GPT-5.6’s 272K price line, Gemini’s shorter output ceiling, and Claude’s finer tokenizer all turn “the same 1M” into different bills and different truncation points. Measure on your regression set, not the sticker alone.
Do we still need RAG if we have a 1M window?
Yes. 1M is for the pack that must be visible together. When the corpus is far larger than the window, or only a small span is relevant each time, retrieval is cheaper, easier to authorize, and easier to refresh. Many production systems retrieve candidates, then read them in a long window — not one or the other.
Why does the model still drop fields after I paste the whole JSON repo?
The window makes things visible. It does not make the model speak a contract. Missing keys, type drift, and markdown fences are shape problems — use Structured Output / Schema and validate again locally. If material in the middle is diluted, effective recall drops too. Only legal, same-origin, already-diffed JSON should enter the window.
Summary and next steps
1M tokens is the 2026 flagship scale: one request can see about a million tokens at once. That makes a whole repo, a whole handbook, or a multimodal pack feasible. It does not automatically mean stronger reasoning, longer output, or a lower unit price. GPT, Claude, and Gemini fork on output caps, surcharge lines, and modalities — not on a single “has 1M” checkbox.
Next: count your handbook and repo with the official tokenizer; reserve output and headroom using the budget above; confirm surcharge thresholds on the three model cards. JSON that will enter the window should go through parse, Schema, and Diff in JSONVue first. How to pin shape: Structured Output. How to layer failures: the JSON errors guide.