Tutorial

Why AI agents are moving from Tool Calling to Skills + Plugins: 2026 architecture change and the JSON data flow

Tool Calling is not obsolete. What is obsolete is stuffing every tool schema and the whole handbook into one request.

We already covered what the agent loop is, how this MCP hop validates arguments, and how to choose Skills / MCP / Plugin. This post does not retell those three. The line you hear in review is different: “We already have Tool Calling — why Skills and Plugins?” That is an architecture question, not a field-filling question. The reasons are concrete. Too many tools, and the full table of schemas eats the window. Too long a process, and the SOP in the system prompt is truncated. Switch clients, and the brief and the hand fork. Agent Skills shrinks discovery to name + description via progressive disclosure. Agent Plugins 1.0.0 specifies the box, not the call. Google says it too: a single skill, a single MCP, a single client does not need a box. This piece lays the old request next to the new trace as JSON.

The call is fine; the discovery surface is not

Tool Calling is still the hop where the model picks a function name and arguments. Nobody retired it in 2026. What retired is a packing habit: at startup, pour fifty inputSchema documents into tools[], then paste “check the project first, then billing, write the weekly summary the finance way” into system. Ten tools are fine. A third business domain later, the request itself becomes a giant JSON nobody diffs. A wrong tool looks like hallucination. The usual root is a discovery surface that is too wide and a brief that got cut. The call hop did not break. You asked it to do discovery as well.

The window is only the first cut. The second is change. Finance reorders the weekly-summary sections, and you edit the system prompt, an internal wiki, and three IDE pastes. Without a diffable SKILL.md, the change never enters a PR. The third cut is cross-client: Cursor’s MCP dialect, Claude Code’s skill directory, Antigravity’s plugin layout — three wrappers. The tools are the same; the boxes fork first. The pain is not the tools/call envelope. It is the layer outside the envelope: when to use it, and what must travel with it.

So “moving to Skills + Plugins” is not a slogan swap. It pulls discovery and packaging out of the call request. The call hop stays — see MCP and JSON Schema. The agent loop stays — see the Agent definition post. Today we only watch how JSON flows after the split. What the client holds after install is Plugin Manifest in practice — five hops. This post is why those five hops appeared.

This layer Old habit 2026 habit
When to use a processA long paragraph in the system promptSkill description (~100 tokens)
Which hands existThe full tools[] in the requesttools/list after a Plugin install or MCP connect
The actual callModel arguments → runtimeUnchanged: still parse + Schema

Not a replacement: the call layer stays, a layer sits on top

Call Skills “the next Tool Calling” and you will debug the wrong layer in the logs. A Skill has no tools/call. It is a brief. Startup injects metadata only. The model matches a trigger, then reads the body and scripts/. MCP is the hand. A Plugin is a directory. The three sit above the call; they do not replace the arguments hop. A honest trace still shows one legal tools/call. If it does not, discovery blocked the model — the call layer was not deleted.

Progressive disclosure is the core of the Agent Skills spec, not a marketing adjective. name max 64, description max 1024, third person, what it does and when to use it. An internal codename or a first-person slogan zeroes the discovery surface — the tools are there, the model never picks them. Loading the body on demand keeps the runbook from fighting fifty schemas for the same window. Scripts are still argv, not first-class tools from tools/list. Stable JSON input still means MCP.

A Plugin is thinner. It cannot inline a tool list, and it cannot put an SOP on the top of plugin.json. The box answers whether this brief and this hand are present. The call answers whether this hop’s arguments are legal. Two JSON files, two jobs. The four decision questions live in the choice post; the five hops after install live in the Manifest post. Today we only draw how the flow went from flat to thick.

The old flow: one giant tools array

The old request looks like a menu plus house rules. Every tools item carries a full inputSchema. system holds the process. One user sentence, and the model orders off the whole menu. A short menu is fast. When the menu is invoices + expenses + IAM + deploy + docs search, ordering fails first: similar tools compete, and the SOP is cut after “never commit secrets…”. The ticket says the model went rogue. A diff says the request body went rogue first.

That JSON has a hidden cost: the runtime assembled it, so it often never enters the repo. Someone adds send_slack and no PR shows how much the schemas grew. Review is stuck with that giant body in the logs. No budget, no signal that discovery should be split out. A flat architecture is not a moral failure. It is growth without a meter.

Below is a flattened old fixture. Production is longer. Make it parse, then put a CI budget on tools.length and request bytes. Crossing the budget is not “add another prompt line.” It is time to move the SOP and the full schemas out of the startup request.

{
  "era": "flat-tool-calling",
  "system": "Always check the billing project first. Never commit secrets. Write the weekly invoice summary the finance team actually reads.",
  "tools": [
    {
      "name": "query_invoices",
      "inputSchema": {
        "type": "object",
        "required": ["week"],
        "properties": {
          "week": { "type": "string", "pattern": "^[0-9]{4}-W[0-9]{2}$" },
          "status": { "type": "string", "enum": ["open", "paid", "overdue"] }
        }
      }
    },
    {
      "name": "export_csv",
      "inputSchema": {
        "type": "object",
        "required": ["week"],
        "properties": { "week": { "type": "string" } }
      }
    },
    {
      "name": "send_slack",
      "inputSchema": {
        "type": "object",
        "required": ["channel", "text"],
        "properties": {
          "channel": { "type": "string" },
          "text": { "type": "string" }
        }
      }
    }
  ]
}

The new flow: metadata → body → tool contract

The new trace thickens on demand. Startup: only skill metadata in context (plus plugin identity, if a box is installed). Match: then read the SKILL.md body. Need a hand: then put this round’s inputSchema from tools/list into the window. Call: arguments still pass the same Schema check. What sits in the window at once goes from “full menu + full house rules” to “an index + the book you opened.” Tokens move from a startup bill to a match bill.

Notice tools[] did not vanish. It moved later. The cost of moving later is one more handshake and one more way to fail the match. A weak description and the model never reaches hop three; users say “we installed a Plugin and it still can’t.” That is a discovery bug, not Tool Calling coming back. The inverse — a trace that already has inputSchema and still pours the runbook into system — is the new architecture sliding backward. The window bill returns.

Below is the new fixture for the same task. It is not a spec file. It is a review trace. Diff it against the old request: the SOP moved from system to skills[].description; the schemas moved from startup to afterMatch.tools. The query_invoices contract on the call layer should stay one canonical Schema. Do not invent a second field set “for the new architecture.”

{
  "era": "skills-plus-plugins",
  "startup": {
    "plugin": "invoice-ops",
    "skills": [
      {
        "name": "write-weekly-summary",
        "description": "Turn invoice query results into the weekly summary finance reads. Use when the user asks for a week-end report.",
        "loaded": "metadata"
      }
    ]
  },
  "afterMatch": {
    "skillBodyLoaded": true,
    "tools": [
      {
        "name": "query_invoices",
        "source": "mcp:invoice-tools",
        "inputSchema": {
          "type": "object",
          "required": ["week"],
          "properties": {
            "week": { "type": "string", "pattern": "^[0-9]{4}-W[0-9]{2}$" },
            "status": { "type": "string", "enum": ["open", "paid", "overdue"] }
          }
        }
      }
    ]
  }
}

Why a Plugin still: the brief and the hand must travel together

Turning the SOP into a Skill and the API into MCP already thins the request. A second client, and directory layout plus MCP dialect start to fork. A Plugin earns its keep when the brief and the hand must travel together. The Google Cloud Developer Plugin packs the gcloud guardrail skill with Developer Knowledge MCP for that reason — not because Tool Calling was too weak. The box executes no tool. It only keeps the discovery path from forking across Antigravity, Claude Code, and Cursor.

One client, one server: keep native MCP config. Shipping a Plugin “to be modern” is just another Manifest to maintain. Questions three and four in the decision post still apply. The architecture shift is not “everyone upgrades to Plugin.” It is “keep the call layer stable; add discovery and packaging when the pain shows.” No second client, and the brief does not depend on the hand: stopping at a Skill or native MCP is a correct 2026 architecture.

Draw the A2A line again. How another agent is found is an Agent Card, not stuffing that agent into tools[]. The other way a flat Tool Calling request explodes is treating a remote agent as one function. That bursts the shape of arguments. Horizontal delegation is not today’s flow. Today is only how this one agent carries less menu and more index.

Symptom Move this layer first Do not
200KB request, 40 entries in toolsDiscovery: skill metadata + schemas on demandLengthen the system prompt again
After switching IDEs, brief and hand disagreePackaging: Plugin directoryCopy another client dialect
Right tool, weekly summary still wrongBrief: SKILL.md bodyAdd an empty format_report tool

Diff the two fixtures side by side

Keep at least three review fixtures: the old request above, the new trace, and one real query_invoices arguments object. The first two own “which layer gave up the bloat.” The third proves the call layer was not rewritten — inputSchema is still the canonical one. Someone copying the Plugin name into tool arguments, or inputSchema into plugin.json, shows up in a diff immediately.

Add two negatives: a new trace with empty skills that still claims era is skills-plus-plugins; an old request over budget on tools.length with no split record. The first catches a slogan change without a discovery change. The second catches “just make the prompt longer.” No secret belongs in a committed fixture.

You can do this in the browser:JSON formatterto see whether the old request and the new trace parse;JSON Schema validatorto check inputSchema and arguments;JSON Diffto put old system / tools next to new skills / afterMatch. Nothing leaves the machine. Further reading:MCP and JSON Schema, the decision post, and the five Manifest hops.

Related: What is an AI Agent, MCP and JSON Schema, Skills vs MCP vs Plugins, Plugin Manifest in practice.

FAQ

Should we delete Tool Calling?

No. The model still picks a function name and arguments; the runtime still issues tools/call. What you delete is the full menu at startup, not the call hop. If the trace has no legal call, check discovery first, then the runtime.

Skill only, no Plugin — is that the new architecture?

That is the discovery layer. The SOP left the system prompt and loads on demand; the request gets thinner. It is not the packaging layer. One client is enough: stop there. A second client appears, and the brief must travel with the hand: then add a Plugin.

Isn’t leaving every schema in the startup request faster?

When the menu is short, the tools are stable, and you have one client — yes, and simpler. Past the budget, in count or in bytes, wrong tool picks eat the latency you saved. Set the budget before you trust the feeling.

Does this fight Programmatic Tool Calling?

No. Programmatic Tool Calling is structure on the call layer: the model chains tools on purpose. Skills + Plugins are discovery and packaging: which brief to read, which hand to connect. Index first, then chain.

Takeaways and next steps

In 2026, “from Tool Calling to Skills + Plugins” collapses to one line: the call hop stays; discovery and packaging leave the giant request. The model still orders. The whole menu is no longer on the table at seat-down.

Ship in this order: budget tools.length and bytes on the old request; fold the SOP into a description; move full schemas to after the match; add plugin.json when a second client appears. Put the old request and the new trace side by side in JSONVue. The loop: the Agent post. Arguments: the MCP post. Whether to pack: the decision post.