All posts
MCP education··8 min read

How does AI function calling and tool execution work?

Function calling is the model emitting a JSON object that names a function and its arguments. Tool execution is your code running that function, appending the result to the message list as a tool message, and calling the model again so it can read what happened. The loop repeats until the model answers with prose instead of a call, which means two round trips minimum for one tool and more for multi-step work. Here is the loop in 20 lines, and the five things that broke ours in production: streamed calls colliding on one index, signed provider extras that must ride back untouched, calls written as prose, models claiming writes they never performed, and a tool list that costs accuracy as it grows.

⌘
Orhan
Founder, Command+K

Function calling is not the model doing anything. The model receives a list of tool schemas alongside the conversation, and when it wants one, it stops generating prose and emits a JSON object naming the function plus its arguments. That is the entire model-side story. Everything after it is yours: you parse the JSON, you execute the real operation, you append the result to the message list as a tool message, and you call the model again so it can read what happened. The loop runs until the model replies with prose instead of a call. I build this loop for a living, so below is the mechanic first, then the five things that broke it in production.

The loop, minus the framework

Strip away every SDK and one tool-calling turn is a while loop over a growing array of messages. This is the shape of it:

ts
const messages = [
  { role: "system", content: systemPrompt },
  { role: "user", content: "close the billing ticket from yesterday" },
];

for (let round = 0; round < 8; round++) {
  const res = await model.chat({ messages, tools });
  const msg = res.choices[0].message;
  messages.push(msg);                       // the request, verbatim

  if (!msg.tool_calls?.length) break;       // prose: the turn is over

  for (const call of msg.tool_calls) {
    const args = JSON.parse(call.function.arguments);
    const result = await runTool(call.function.name, args);
    messages.push({
      role: "tool",
      tool_call_id: call.id,                // must match, or the call dangles
      content: JSON.stringify(result),
    });
  }
}

Four details in that snippet are load bearing. The assistant message goes back into the array untouched, because the next request has to contain the call the model made. Every tool_call_id needs a matching tool message before the next request, or the provider rejects the array. The loop is bounded, because a model that keeps calling tools will keep calling tools. And the result is a string, so whatever your tool returns has to be serialized into something the model can read.

That is the whole mechanic. A tool-calling agent is this loop with a good tool list and a system prompt. What makes it hard in production is that every piece of it has a failure mode the tutorials skip.

1. Streaming does not hand you a tool call

Non-streamed responses give you tool_calls whole. Streamed responses give you pieces. A single call arrives as a fragment carrying the id and the function name, then a series of fragments carrying slices of the argument string, each tagged with an index. You group by index and concatenate:

json
{"index":0,"id":"call_1","function":{"name":"knowledge_search","arguments":""}}
{"index":0,"function":{"arguments":"{\"query\":"}}
{"index":0,"function":{"arguments":"\"pricing\",\"limit\":5}"}}

Our bug was the case that is not in the docs. A provider sent a second, already complete call at index 0 after finishing the first one. Our merge appended it to the existing accumulator, so the arguments field became two JSON objects glued together and JSON.parse threw. From the outside the widget looked like the model had gone stupid mid-turn. The fix is to inspect each fragment: if it carries an id or a function name while the slot already holds a named call, it is a new call, so open a new accumulator rather than appending to the old one.

2. Provider extras have to ride back untouched

Gemini attaches a thought_signature to tool calls, inside extra_content.google. It is a signed token over the reasoning that produced the call, and the next request has to carry it back on the assistant message, byte for byte. Drop it and the provider rejects the continuation. Synthesize a tool call yourself, without a signature, and it is rejected too.

You cannot fake an assistant tool call

This is worth internalizing before you design anything clever. Injecting a fabricated tool_calls entry into history (to replay an approved action, or to stitch two turns together) works on some providers and fails hard on signed ones. Replay the real message you stored, extras included, or re-plan the turn from scratch.

Practically: when you persist a turn, persist the provider's assistant message as it arrived, not a normalized version of it. We learned that by normalizing ours and having every Gemini continuation rejected.

3. Sometimes the model writes the call as prose

Every so often a model does not emit a tool call at all. It writes what the call would look like, in the text channel, as markup. No tool_calls array arrives, so your loop sees prose, breaks, and ships raw markup to the user.

We handle it in two steps. If a round produces text that matches tool-call markup, we parse it and synthesize the call ourselves for the small set of internal tools where that is safe, which recovers the turn. Otherwise we strip the markup out of the visible answer and force a prose-only retry with tools disabled, so the user gets a real sentence instead of a leaked function signature. Log every time this happens with the model name: it is the cheapest signal you have that a prompt or a model swap has degraded tool discipline.

4. The model will claim it did the write

This is the failure that actually costs trust. The user asks for an action. The model writes "Done, I've archived the post" and never emits a single tool call. Nothing happened. The user believes it did.

Our guard starts as a streaming decision. When the turn is action-shaped and no tool has executed yet, we hold the text deltas instead of streaming them out. If a tool call shows up, the hold releases and the text flows. If the round ends with a write claim and no execution, we push a system nudge and give the model one more round to actually call the tool. If it still refuses, we replace the claim with an honest message before anything reaches the user.

ts
const writeGuardArmed = writeContext && executedToolCalls === 0;
const gate = createDeltaGate({ hold: writeGuardArmed, send: (t) => send("delta", { delta: t }) });
// gate.toolCallSeen() releases the hold the moment a real call arrives

You cannot fix this with prompting alone. I tried, for weeks. The instruction "never claim an action you did not perform" helps the average case and does nothing for the tail, and the tail is where a customer sees a lie. The runtime check is the fix, and it belongs on the stream so the lie never gets printed.

5. The tool list is a cost and an accuracy problem

Every schema you pass sits in the prompt on every round of every turn. A deployment with 84 tools pays for all 84 on all 8 rounds, and accuracy drops as the list grows, because near duplicates give the model more ways to pick wrong. Once a deployment passes 24 tools we stop sending the catalog and retrieve a subset per turn instead: the top 20 by relevance to the user message, plus whatever the turn has already used. Prompt cost fell, and tool selection got better rather than worse.

The other half of accuracy is the description text. The model chooses purely from names, descriptions and parameter docs, so a tool described as "updates a resource" is a coin flip against three sibling tools. Write descriptions for a reader who cannot see your API.

Order of operations, written down

  1. Send the conversation plus the tool schemas.
  2. Read tool_calls. If absent, the turn is prose and you are done.
  3. Push the assistant message back verbatim, provider extras and all.
  4. Execute each call. Return failures as tool results, not exceptions.
  5. Push one tool message per tool_call_id, then loop, up to your round cap.
  6. On the final round, drop the tools so the model has to answer in prose.

That last one matters more than it sounds. A model given tools on its last permitted round will often emit a call you cannot execute, and then your turn has no answer in it at all.

Where MCP fits

Nothing above changes when you move to MCP. The loop is the same loop. MCP replaces the hardcoded tool array with a list you fetch from a server at runtime, and runTool becomes a call over the wire instead of a local switch. The win is that the same tool definition works in Claude, Cursor, ChatGPT and your own widget, so you write the schema once. The cost is a network hop and a new set of failure modes at the transport layer, which we wrote up separately after taking a real deployment through production clients.

If you want the loop without building it, that is what we sell: point Command+K at your existing API and the tool list, execution, approvals and round capping come with it. Either way, build the guard in step 4. The day a model tells your customer it refunded an invoice it never touched is the day you wish you had.

FAQ

Does the model execute the function itself?
No. The model only emits a JSON object naming a function and its arguments. Your code executes it, and your code appends the result back into the message list as a tool message. If you never run anything, the model has no idea: it will happily write 'Done, I created the ticket' with nothing behind it.
How many round trips does one tool-calling turn take?
At least two: one for the model to request the call, one for it to read the result and answer. Multi-step work takes more, because each new tool call needs its own round trip. We cap our widget at 8 rounds per turn when tools are available, and 1 when they are not.
Why do my streamed tool calls arrive with broken JSON arguments?
Streamed tool calls arrive as fragments keyed by an index, and you concatenate the argument strings per index. If a provider reuses an index for a second, complete call in the same turn, a naive merge glues two JSON objects together and JSON.parse throws. Detect a fragment that starts a new call (it carries a fresh id or name) and open a new accumulator instead of appending.
What is the difference between function calling and MCP?
Function calling is the per-model mechanic described here: schema in, JSON call out. MCP is the protocol that advertises and transports those tools so one tool definition works across models and clients. The execution loop is identical either way; MCP only changes where the tool list comes from and who runs the call.
Should the model see tool errors?
Yes. Send the failure back as the tool result, with a message the model can act on, rather than throwing inside your loop. A model that reads 'issue_id not found' will go look up the id. A model that gets silence will invent one.

Keep reading

Go deeper