Tutorial

What is an AI Agent? 2026 guide to how agents work, tool calling, function calling, and JSON

A chat window only replies. An agent picks a tool, fills arguments, reads the result, and decides the next step. The spine is not magic — it is JSON: tool defs, arguments, and results.

In 2026 “AI agent” shows up in launches, job posts, and architecture reviews — and rarely means the same thing. A chat window with plugins, a scheduled workflow, an IDE assistant on MCP: all get the label. This article uses an engineering definition: an agent is a model-driven runtime that can loop on tools and hand off state in structured data (almost always JSON). It is not a chattier model. It is model + tool runtime + contract. We split what tool calling, function calling, and JSON each govern, then link our Structured Output, MCP Schema, and A2A guides.

What an AI agent is — vs chat and workflows

The minimum definition needs three things: a goal (what the user wants done), perception (context plus tool receipts), and action (which tool and which arguments, or a final answer). The model chooses the action each step; the runtime executes tools and writes observations back. No loop, no tools, no checkable argument shape — that is just chat.

The difference from a chatbot is the stop condition, not the brand. Chat can end after one turn. An agent should not claim the task is done before tool results arrive — it may search invoices, summarize, then ask you to confirm. The difference from a classic workflow is who draws the edges: n8n / Temporal next hops are wired by people; an agent’s next hop is chosen from the current JSON observation. Workflows are predictable and replayable; agents are flexible, and they make “wrong arguments” a first-class failure.

Common 2026 shapes: coding agents (read files, run tests, patch), support and ops agents (orders, tickets), and multi-agent orchestration (a planner delegates to specialists). The contract is the same: tool boundaries in JSON Schema; arguments and results are parseable JSON. Apple can expose one function to App Intents and model tools — see Apple AI Agent and JSON.

How it works in 2026: observe → decide → call → observe

Peel the demo video and a typical agent loop is five steps:

  1. The user goal enters context (natural language plus optional system constraints).
  2. The runtime injects the tool list: name, description, and a JSON Schema (parameters / inputSchema).
  3. The model returns tool_calls (which function, which arguments) or a final text / Structured Output.
  4. The runtime JSON.parse arguments, runs a local function, HTTP, or MCP tools/call, and writes result JSON back into the messages.
  5. The model reads the result and picks the next tool or stops. maxSteps, user cancel, or Schema failure also stop the loop.

The loop shape itself can be config, not framework magic. The JSON below describes hops and stop conditions only — business fields belong on each tool’s own Schema.

{
  "loop": "agent",
  "maxSteps": 8,
  "stopWhen": ["final_answer", "max_steps", "user_cancel", "schema_fail"],
  "hops": [
    { "kind": "model", "emits": "tool_calls | text" },
    { "kind": "runtime", "emits": "tool_result JSON" },
    { "kind": "model", "emits": "next_tool | final JSON" }
  ]
}

Failures cluster at steps 3→4: arguments are a string treated as an object, numbers as strings, missing required, or a tool name that drifted from a cached list. A “smarter” model does not fix contract drift. Stateless MCP attaches protocol metadata per request; argument shape still depends on the Schema you fed the model.

Tool calling vs function calling: two names, one mechanism

Engineering-wise they are the same: the model does not touch the database; it emits a structured “please call this function with these arguments” request, and the runtime executes it. Product names shifted from 2023 to 2026:

Concept Function calling Tool calling
Origin OpenAI from 2023: function_call / functions[] Industry umbrella 2024–2026 across vendor APIs and MCP
What the model emits function.name + arguments (usually a JSON string) OpenAI tools[], Anthropic tool_use, Gemini functionCall
Where Schema lives function.parameters tools[].function.parameters or MCP inputSchema
Vs Structured Output Does not govern the final answer — only this hop’s inputs Same split: tool hop and final-answer hop must be separate files

OpenAI later folded functions into tools and added strict. Anthropic uses tool_use / input_schema. Gemini uses function declarations. MCP uses JSON-RPC tools/call. Names differ; arguments are still a JSON object. “Function calling is obsolete” is usually a bad migration reason — the old field names aged, not the mechanism.

The same invoice search as an OpenAI strict tool. Under strict, every property must appear in required, or the model may legally omit fields you assumed defaulted:

{
  "type": "function",
  "function": {
    "name": "searchInvoices",
    "description": "Search invoices by date range and status",
    "strict": true,
    "parameters": {
      "type": "object",
      "properties": {
        "startDate": { "type": "string", "format": "date" },
        "endDate": { "type": "string", "format": "date" },
        "status": {
          "type": "string",
          "enum": ["draft", "sent", "paid", "void"]
        }
      },
      "required": ["startDate", "endDate", "status"],
      "additionalProperties": false
    }
  }
}

When the model calls it, a typical tool_calls item looks like this. arguments is still a string: parse, then validate. A failed JSON.parse is not automatically “the model is broken” — separate syntax failure from Schema mismatch. Taxonomy: AI JSON errors guide.

{
  "id": "call_8f3a",
  "type": "function",
  "function": {
    "name": "searchInvoices",
    "arguments": "{\"startDate\":\"2026-01-01\",\"endDate\":\"2026-01-31\",\"status\":\"paid\"}"
  }
}

Why JSON is the agent’s contract language

An agent pipeline carries at least three JSON payloads that should share one canonical Schema:

  1. Tool definition: name / description / parameters (or MCP inputSchema).
  2. Model arguments: the chosen keys, often as a string inside tool_calls.
  3. Tool result: the runtime’s ok / result or error envelope, so the model can choose the next hop.

A fourth payload is optional: Structured Output for the final answer. That file describes the answer to the user or downstream system — not what searchInvoices needs. Same syntax, different semantics; split files and versions. See AI Structured Output tutorial.

After a successful tool run, return a stable envelope instead of dumping raw downstream HTTP at the model. The result below exposes only fields the business needs; the raw response goes to logs. Stable shape makes retries and summaries predictable:

{
  "toolCallId": "call_8f3a",
  "name": "searchInvoices",
  "ok": true,
  "result": {
    "count": 2,
    "items": [
      { "id": "INV-1042", "total": 1280.5, "currency": "USD" },
      { "id": "INV-1048", "total": 640.0, "currency": "USD" }
    ]
  }
}

2026 stacks: OpenAI, Anthropic, Gemini, MCP

There is no portable Schema everywhere, but tool-calling data structures converge: arguments are a JSON object; the contract is a JSON Schema subset.

Stack / protocol How tools are declared How a call is emitted
OpenAI Chat / Responses tools[].function.parameters, optional strict tool_calls[].function.arguments string
Anthropic Messages tools[].input_schema input object on a tool_use block
Gemini function_declarations.parameters functionCall.args; final JSON uses responseSchema
MCP 2026-07-28 Tool.inputSchema (the protocol does not validate for you) params.arguments on JSON-RPC tools/call

MCP is discovery and transport, not a type system. inputSchema declares shape; the server still parses and Schema-validates. Field mapping and one Schema on three surfaces: MCP and JSON Schema. Multi-agent handoffs use A2A message/send; skill lists hang Schema too — that is agent-to-agent, not model-to-tool. Compare A2A vs MCP.

Stateless MCP scales horizontally after sessions leave the protocol; it does not auto-fix arguments. A stale tools/list cache writes the wrong shape into the call. Protocol details: Stateless MCP guide.

Validation on the ground and JSONVue

Every hop: parse → Schema → business rules. A clever model does not replace those three steps.

  1. arguments string: JSON.parse; on failure log raw + tool_call id and return a retryable error envelope.
  2. Validate against parameters / inputSchema (Draft 2020-12); emit path and keyword.
  3. Business gate: date range, enum vs permissions, foreign keys. Only then hit the downstream API.

Line up three blobs: model arguments, the body you send to MCP or HTTP, the object the server actually used. Mismatch is almost always the adapter. In the browser: JSON formatter for parse; JSON Schema validator for arguments vs the Schema file; JSON Diff for model arguments vs the downstream body. Share valid / missing-field / wrong-enum fixtures in CI and manual debug.

Further reading: MCP and JSON Schema, Structured Output, AI JSON errors guide, A2A vs MCP.

FAQ

Are agents the same as RAG?

No. RAG stuffs retrieved documents into context so the model can answer; it can be one tool (searchDocs) inside an agent. Without a tool loop, RAG is still augmented Q&A.

Is function calling obsolete — should we only say tool calling?

Docs and SDK titles moved; the mechanism did not. OpenAI still uses type: function tools; Anthropic, Gemini, and MCP use their own field names. Keep one canonical Schema and generate vendor shells. Do not rewrite the business because a heading changed.

We enabled Structured Output — do we still validate tool arguments?

Yes. Structured Output constrains the final answer; arguments are another hop. The usual incident is “answer Schema passed, tools/call still missing keys.” Split files; parse + validate every hop.

Can we build an agent without MCP?

Yes. MCP is one protocol for discovery and remote calls, not the definition of an agent. Local functions, OpenAPI, and your own HTTP all work if arguments and results are validatable JSON. MCP’s value is a standard catalog and transport — especially remote and multi-client.

Summary and next steps

A 2026 AI agent: the model chooses actions in a loop, the runtime executes tools, JSON Schema describes each hop. Tool calling and function calling are product names for one mechanism. MCP, A2A, and Structured Output govern different hops — do not force one file onto all of them.

Next: list the three JSON payloads in your system (definition, arguments, result) and check they share a Schema; replay valid / missing-field / wrong-enum in JSONVue. Protocol details in the MCP article; final-answer shape in Structured Output; failure layers in the JSON errors guide.