All posts
MCP education··7 min read

Capping each tool result did nothing: an agent turn needs one budget

Our assistant kept stopping mid-sentence on the most expensive turns. The cause was a per-result truncation cap that said nothing about the total: seven results, each legally under 8,000 characters, carried about 56,000 characters of evidence into the final round. Here is the shared 24,000-character turn budget that replaced it, why it fills like water instead of splitting evenly, why the trimmed copy must never be the logged copy, and the second half of the bill nobody meters, which is the tool list resent on every round.

⌘
Orhan
Founder

A customer told me the assistant kept stopping mid-sentence. The turns were also our most expensive. I assumed those were two bugs. They were one, and the cause was a truncation rule I had written myself and considered done.

The symptom was a cut answer, not a bill

The report was about an assistant connected to a knowledge base, asked something like "which of my published posts has this feature not been mentioned in yet?". The answer began well and then stopped, in the middle of a word. Our widget marks that case honestly, so the visitor saw a notice that the reply hit its length limit.

The obvious reading is that the output cap is too tight. Ours is DEFAULT_OUTPUT_TOKENS = 3072, with a hard ceiling of 8,192. Raising it would have made the symptom disappear and the bill worse, which is how you know it was not the fix.

Every result passed its own cap

We truncate tool results before they go to the model. The line is about as simple as it looks:

ts
const resultStr = JSON.stringify(result).slice(0, 8000);

Eight thousand characters is a generous but sane cap for one result. The question I had never asked is how many results arrive in a single turn. Our loop allows up to eight rounds when tools are connected, and the model picks what to call in each one. The question above made it read seven full article bodies.

Seven results, each one legally under the cap, each one roughly 8,000 characters. The final round carried about 56,000 characters of evidence into the request. That is what made the answer long enough to hit the output ceiling. The cut answer was a downstream symptom of an input that nobody had bounded.

A per-item cap is not a budget

A per-result cap limits the worst single result. It says nothing about the total, because the number of results in a turn is a decision the model makes at runtime. Any limit you want to hold has to be expressed over the turn.

One budget for the turn

Right before the final round, the turn's tool messages are now trimmed together to a shared budget:

ts
export const TOOL_RESULT_TURN_BUDGET_CHARS = 24_000;
export const TOOL_RESULT_MIN_KEEP_CHARS = 1_200;
export const TOOL_RESULT_TRIM_MARKER =
  "\n… (trimmed to fit this turn's evidence budget — ask for one item if you need its full body)";

Three things in there matter more than the number 24,000, which you should tune to your own model and prompt.

  • The floor. Nothing is cut below 1,200 characters. A result trimmed to 200 characters is not cheaper evidence, it is noise with a token cost.
  • The marker. Every trimmed result ends with a visible line saying so. Without it the model reads a result that stops cleanly and treats it as the whole thing, which is how a truncation turns into a confident wrong answer.
  • The timing. The trim runs before the final round, not as results arrive. Mid-loop the model may still need the full body to decide what to call next.
  • The scope. Only tool messages are touched. The system prompt and the user turns are somebody else's budget.

Why dividing the budget by the number of results is wrong

The first version I wrote gave every result an equal share. With a 24,000 budget and seven results that is 3,428 characters each. It is one line of code and it wastes most of the budget.

A real turn is lopsided. Say the model called one tool that returned a 300-character issue count, one that returned a 420-character status, one 800-character list, and four full article bodies at the 8,000 cap. That is 33,520 characters for a 24,000 budget. Equal shares would cut the four bodies from 8,000 to 3,428 and leave the three small results untouched at their own sizes, for a total of 15,232. You would have thrown away 8,768 characters of budget you were willing to pay for, taken out of the only results that contained the answer.

So the trim fills like water instead. Results are sorted by size. Each one is offered an equal share of what is left; anything that already fits under its share is kept whole and releases its slack to the rest. The cap settles where the remaining large results exhaust what is left.

ts
const sorted = [...lens].sort((a, b) => a - b);
let remaining = budgetChars;
let cap = minKeepChars;
for (let k = 0; k < sorted.length; k++) {
  const share = Math.floor(remaining / (sorted.length - k));
  if (sorted[k] <= share) {
    remaining -= sorted[k];
    continue;
  }
  cap = share;
  break;
}

On the lopsided turn above the cap lands at 5,620 instead of 3,428. The three small results survive whole, the four bodies keep 5,620 characters each, and the turn uses 24,000 of its 24,000. Same ceiling, 2,192 more characters of real evidence in each of the results that mattered.

When every result is large the two rules converge, which is the point. Seven results at 8,000 characters give a cap of 3,428 and a total of 23,996. Water-filling only pays off when the turn is uneven, and real turns almost always are.

The model's copy is not the log

One detail I would get wrong if I rewrote this from memory: the trim mutates only the message array being sent to the model. The conversation row in the database and the tool card the visitor expands in the widget both keep the full result.

If you trim the stored copy you have quietly made your audit trail a function of your token budget. Six weeks later someone asks what a tool actually returned on a disputed write, and the honest answer is "the first 5,620 characters of it".

The other end of the prompt is the tool list

Results are the half of the request that grows during a turn. The tool list is the half that is already large before the turn starts, and it is sent again on every round.

One of our deployments exposes 84 GraphQL tools. Eight rounds means 84 JSON schemas sent eight times for a turn in which the model calls two of them. So over a threshold we select per turn instead of sending everything:

  • Under 24 candidate tools, send them all. Retrieval costs a round trip and is not worth it on a small connector.
  • Over it, embed the tool names and descriptions and select the top 20 for this turn's query, inside a 4-second budget. If retrieval is slow or throws, the turn falls back to the full list rather than failing.
  • A must-keep set overrides the ranking: identity tools such as me or get_current_user, anything pinned by the binding's use-case policy, and every tool already called in the last 12 tool messages of this conversation. A follow-up question must not lose the tool that answered the first one.

Tool selection is where tool descriptions stop being documentation and start being retrieval keys. A tool whose description is "Updates the record" will not be selected for anything, and you will read it as the model being bad at tool choice.

What to check in your own agent loop

  • Log the total characters of tool evidence per turn, not per call. If you only have the per-call number you cannot see this bug at all.
  • Count how many times your tool list is serialized in one user turn. Multiply by your rounds. That is usually the largest single line in an agent bill.
  • Make every truncation visible to the model. A silent cut is worse than a short result.
  • Check whether raising your output cap would fix a cut answer or just make it longer. Ours was the second one.

We run this inside the Command+K widget because our own customers connect connectors we did not write and cannot predict the size of. If you want the rest of the cost surface rather than this one slice, the cost optimization tactics post covers the other levers.

FAQ

Why is my agent's answer getting cut off mid-sentence?
Usually because the evidence you fed it was too long, not because the output cap is too low. A model handed 56,000 characters of article text writes a long answer, and a long answer runs into the output token limit. Cap the evidence for the whole turn before you raise the output cap, and the answers get shorter on their own.
Should I truncate each tool result or the whole turn?
The whole turn. A per-result cap tells you nothing about the total, because the number of calls in a turn is chosen by the model, not by you. Seven results that each pass an 8,000-character cap still ship 56,000 characters. One shared budget is the only limit that holds no matter how many tools the model decides to call.
Does sending 80 MCP tools to the model cost anything if the model only calls two?
Yes, and you pay it on every round. The tool list is part of the request, so an agent loop that runs eight rounds sends all 80 JSON schemas eight times. Selecting a subset per turn is the cheapest change available on most agent loops.
How do I truncate a tool result without the model inventing the missing part?
Put a visible marker at the cut. A result that simply stops looks complete to the model, and it will answer as if it saw everything. A trailing line saying the result was trimmed and how to get the full body turns a silent gap into something the model can report or fetch.

Keep reading

Go deeper