Picture handing someone a job and a manual.
The job is small: fix one thing, in one file. The manual is 185 pages long and describes every tool in the building. The lathe. The label printer. The payroll system. The social media scheduler. The conference room booking tool. You hand over both, and then, because of how the arrangement happens to work, you read the entire manual aloud again before every single sentence you say to them.
Nobody would design that. I ran it for months without noticing.
I run a coding agent on my own hardware: two RTX 5090s on a workstation under the desk, driving Claude Code against a locally served model through a router that translates between the two API shapes. I have written about tuning the box, about making its inference server event-triggered rather than burning two CPU cores doing nothing, about the engine bake-off that prefix caching decided, about the attention architecture that made every turn re-read 162,000 tokens, about what actually fits in 64 GB of VRAM, and about the quantization formats that fail while looking healthy.
This one is cheaper to fix than any of those, and I suspect more people have it. It cost nothing to diagnose. It took one afternoon to fix. And the fix was worth roughly half my context window.
The measurements below are from my own box on 3 June 2026. The external research I cite is dated where it matters, because this is a fast-moving area and a stale number here would be worse than no number.
The two numbers that could not both be right
The whole thing started with a contradiction sitting in my own notes.
One line said the injected system prompt for an engineering task was roughly 3,500 to 4,500 tokens. I had counted it. It is mostly instructions about how to behave, which files count as authoritative, and how to close out work. A few pages of prose.
Another line, written weeks later in the same document, said the base prompt was about 99,000 tokens. I knew that one was real too, because I had raised the inference server's context limit to the model's native ceiling after a task ran off the end of it.
Two numbers for the same quantity, a factor of twenty-five apart, both written by me, both correct. Reconciling them is the entire point of this post.
The first number counts the prose. The second measures what actually goes over the wire. The gap between them is almost entirely tool definitions: the JSON schema for every tool the agent is permitted to call, serialised into the request, in full, on every single turn.
I had roughly 185 of them, across about a dozen connected servers, nearly all of them MCP servers. Measured from my router's request log, those definitions alone came to around 55,000 tokens.
This is the tool schema token cost, and it is structurally invisible: it appears nowhere in your source tree, because you never wrote it. Your MCP servers generate it, your client concatenates it, and the only place it exists in full is the request body.
Here is the base prompt as it actually was, before a single file had been read:
| What is in the base prompt | Measured tokens | Share of the base |
|---|---|---|
| Tool schemas, roughly 185 tools across a dozen servers | ~55,000 | just over half |
| Auto-injected project instruction files | ~7,000 | ~7% |
| The harness prose I believed was the whole thing | ~4,000 | ~4% |
| Everything else: the deferred-tool listing and the harness's own scaffolding | the remainder | ~33% |
| Total, measured on the wire | ~99,000 |

Be suspicious of that last row, and I will be honest about it: I measured the tool schemas and the injected instruction files directly, and the prose I had counted by hand. The remaining third I can attribute by category but did not itemise at the time. The shape holds regardless, and the shape is the finding. The part I had counted was four per cent of the thing I thought I had counted.
That is the reconciliation, and once you have seen it you cannot unsee it. When people talk about their agent's system prompt, they almost always mean the prose, because the prose is the part they wrote and the part they can see in a file. The schemas are generated, invisible, and an order of magnitude larger.
Where the window actually goes on a long agent run
Token cost is the boring half. The interesting half is what the base prompt evicts.
Your agent context window budget is a fixed pie with three slices, and only one of them does any work.
My server runs with a hard context ceiling of about 202,000 tokens, which is the model's native maximum. Of that, I reserve roughly 32,000 for output, because the model needs somewhere to write. So the arithmetic for any given task is:
202,000 context ceiling
-32,000 reserved for output
-99,000 base prompt, paid on every turn
--------
71,000 left for the actual work
Seventy-one thousand tokens sounds generous until you watch an agent use it. A coding task is a loop of thirty to sixty turns, and the transcript only grows: every file read, every diff, every test output, every tool result stays in the window. A genuinely medium-sized task blows through 71,000 tokens somewhere in the middle, and I know because that is exactly what happened. I did not predict this and then confirm it. A task ran off the end of the window, and raising the ceiling is what sent me looking for where the window had gone.

Worth saying plainly, because it is the part that surprised me: the base prompt does not compete with your task for the context window. It wins, permanently, before your task starts. The transcript is what gets squeezed. And the transcript is the only part of the window that contains anything about the problem you are trying to solve.
What 185 MCP tools actually cost, and why tool count is the wrong unit
When I tell people this story, the first question is always "how many tools is too many?" That question has a much better answer than my own numbers give it, so here is the outside evidence.
Anthropic published a figure that maps almost exactly onto mine. In their write-up of on-demand tool search, they give a worked example of a five-server setup: "58 tools consuming approximately 55K tokens before the conversation even starts," and note that internally, "tool definitions consume 134K tokens before optimization" (Introducing advanced tool use on the Claude Developer Platform, 24 November 2025).
Look at that against mine. My 185 tools cost about 55,000 tokens. Their 58 tools cost about 55,000 tokens. Same bill, a third of the count.
That comparison is the most useful thing in this post, so I want to be careful with it. It does not mean their tools were wasteful or mine were lean. It means tool count is a proxy, and a bad one. What you pay for is schema bytes: the length of each description, the depth of the parameter object, how many enum members you listed, whether every field carries a sentence of explanation. A dozen chatty tools with deeply nested inputs will cost you more than fifty terse ones. If you are budgeting by counting tools in a config file, you are budgeting in the wrong unit.
The academic end of this gives the same answer at larger scale. MCPVerse, a benchmark built on more than 550 real executable tools, reports that "combined schemas of these tools exceed 140,000 tokens" (arXiv:2508.16260). That is a benchmark whose action space alone would not fit in most production context windows.
And this is exactly why code-execution approaches have taken off: Anthropic's own code execution with MCP post (4 November 2025) reports a single workflow going from 150,000 tokens to 2,000, a saving of 98.7%, by letting the model call tools as a code API it reads on demand rather than loading every definition upfront. That is a bigger lever than the one I pulled. It is also a much larger change, and I will come back to why I took the smaller one first.
How many tools can an LLM agent actually handle
Here is the half that matters more than the tokens, and the half I did not expect to care about.
Choosing the right tool from a list of 185 is close to the worst-shaped problem you can hand a language model. It is a large discrete choice over near-duplicate options, where several candidates are plausible and only one is right. Models are not uniformly good at this, and they get worse in a specific, documented way as the list grows.
Speakeasy ran a controlled version of this that is worth reading because it holds everything constant except tool count. Using a small model (Qwen 3.1 1.7B) against the same task, they report: at 10 tools, the model "successfully retrieved images of four different dog breeds with correct tool names and no errors"; at 20 tools, "19 out of 20 tool calls correct"; at 40 tools, 3 of 4, with a hallucination; and at 107 tools, it "struggled to select a tool most of the time" and "hallucinated incorrect tool names" (Why less is more for MCP).
Note what that is not. It is not a gentle slope. It is fine, fine, wobbling, then off a cliff. And note what it is: one small model. Secondary write-ups of that experiment often report it as large models failing too, which is not what the source says. I checked, because the difference matters to my case.
It matters because my case is a small model. The thing on my desk is a 30-billion-parameter mixture-of-experts with roughly 3 billion parameters active per token. It reasons well for its size, which is why it is there. But "3 billion active" is precisely the regime where a 185-way choice degrades first.
Anthropic's numbers point the same way from the other direction. Moving from all-tools-upfront to on-demand tool search, they measured accuracy on MCP evaluations rising from 49% to 74% for one model and from 79.5% to 88.1% for a stronger one. Those are not token savings. Those are correctness gains from showing the model fewer things.
Now the honest complication, because the picture is not one-directional. MCPVerse found that the strongest model they tested actually did better with a larger exploration space than with an oracle-narrowed one, scoring 61.01% in standard mode against 57.77% in oracle mode. More room to explore helped it. Meanwhile a 2026 paper from Meta on how many candidate tools to surface found the opposite effect for selection accuracy specifically: shorter adaptive candidate lists produced "93.1% versus 87.1% when always shown 5 tools," widening to "76.8% vs 60.9% on medium-difficulty queries" (How Many Tools Should an LLM Agent See? A Chance-Corrected Answer, May 2026).
Reading those together, the rule I take away is not "fewer tools is always better." It is:
The weaker the model relative to the task, the more a wide tool surface costs it. Curation is a correctness intervention aimed at the model you actually have, not a universal law.
If you are running a frontier model on a well-specified task, a big catalogue may genuinely be an asset. If you are running something small, local, or cheap, the catalogue is where it will fail, and it will fail by confidently calling a tool that does not exist.
What I changed: 185 tools down to 24, on one lane only
The fix was a per-lane tool profile. The agent lane that picks up coding tasks now advertises about 24 tools instead of about 185: the board and messaging tools it genuinely uses, plus two small servers it needs for history and for asking a sibling agent a question. Every other consumer of the same infrastructure sees the full surface, byte for byte unchanged.
Mechanically there are two halves, and both are ordinary:
- The tool server learned a profile. One environment variable tells it which named subset to advertise. Set it, and it offers the focused set. Leave it unset, or set it to anything it does not recognise, and it offers everything.
- The agent process gets a smaller server list. When the lane spawns, it builds a trimmed server configuration and starts the agent in strict mode, so the servers not on that list are never loaded at all. This is the half that actually removes the bytes, because a server the process never connects to contributes no schemas.
The design decisions around it mattered more than the mechanism, and they are the transferable part.

Fail open, on every uncertainty. Missing configuration file? Full surface. Unrecognised profile name? Full surface. Anything ambiguous? Full surface. This is an optimisation, and an optimisation that can block work is not an optimisation, it is an outage with a good excuse. The worst thing this change can do when it misfires is fail to save me tokens.
One switch turns the whole path off, with no code revert. The slimming only happens on one routing path, and that routing path already had a single setting that disables it. Flip it and the lane goes back to its previous behaviour. I did not add a second kill switch, because the one that already existed subsumed the new behaviour. If you are adding a context optimisation, know in advance what the one-step undo is, and make sure it is a setting rather than a deploy.
Scope it so a mistake cannot escape. The gate is narrow and explicit: this lane, these two task types, nothing else. Interactive sessions, the cloud path for larger tasks, and the marketing agent were all untouched. That was not caution for its own sake. It is what let me ship the same day rather than scheduling a careful rollout, because the blast radius was small enough to reason about completely.
Verify at runtime, not in review. The check that mattered was not the diff. It was starting the server with the profile set and counting the tools it advertised, then starting it without and counting again: 14 against 74 from that one server, with nothing from the marketing or advertising surfaces leaking into the slim set. A tool surface is an emergent property of a running process. Read it from the running process.
Net effect on the window: roughly 40,000 to 50,000 tokens returned to the transcript. The base prompt stopped being the majority shareholder in the context window.
Why a smaller tool prefix compounds with prefix caching
This is the second-order win, and it is the reason the change was worth more than the token count suggests.
Prefix caching is the feature my whole stack is built around. The engine keeps the computed state for a shared prefix so that turn two does not recompute what turn one already processed. On a roughly 30,000-token shared prefix I measured a cold turn at 19.7 seconds and warm turns at 0.10 seconds, about 190 times faster, at around an 80% hit rate. Without it, an agent loop is unusable on this hardware.
A tool surface is a perfect cache prefix: it is large, it sits at the very front of the request, and it is byte-identical on every turn. So you might reasonably think shrinking it is pointless. It was already cached.
Two things make it worth doing anyway:
- Cache capacity is finite and shared. The key-value cache lives in the same VRAM as the model weights. Every block held for a stable tool prefix is a block not available to hold the growing transcript, which is the part that changes and the part you want retained between turns.
- Caching does not buy back window. A cached prefix is fast, not free. It still occupies its full share of the context ceiling. Prefix caching fixes the time cost of a big base prompt and does nothing at all about the space cost.
So the two techniques address different axes and stack cleanly. Caching makes the prefix cheap to reprocess. Curation makes it small. You want both, and if you only have one, you are still paying the other bill.
What I have not done, and why
The largest remaining consumer of my context window is not tools any more. It is the auto-injected project instruction files plus a set of governance documents the harness reads on the very first turn. Together that is somewhere between 41,000 and 55,000 tokens of first-turn reading, which is now a bigger line item than the schemas ever were.
I know how to cut it. The harness supports starting in a bare mode with a hand-supplied system prompt file, which would let me replace the automatic injection with a short identity prelude and pointers to read the governance documents only when a task actually touches them.
I have not shipped it, and the reason is worth stating plainly: that change alters how the prompt is assembled and when startup hooks fire. Specifically, I do not yet know whether bare mode still fires the session-start hook, whether it still runs the pre-tool-use safety hook that stops the agent doing things I do not want it doing, or how it takes its prompt. One of those is a context optimisation. Another is a safety control. If you cannot tell in advance which of your hooks survive a flag, you do not yet know what that flag does, and a tokens-versus-safety trade made by accident is not a trade, it is a regression.
So it sits behind a probe: establish the hook behaviour empirically first, then decide. Three months on it is still unshipped, and I would rather publish that than imply this was a finished sweep. The cheap half is done. The bigger half is blocked on a question I have not made time to answer.
The checklist: how to measure and reduce agent prompt tokens
If you have an agent with tools wired into it and you have never measured this, here is the order I would follow.
Measure the real base prompt from the wire, not from your source tree. This is the whole thing. Put a proxy in front of your model endpoint, or read your provider's request logs, and look at the actual serialised request for a turn. Count the tokens. Whatever number you had in your head, this one will be larger. If you skip every other item here, do this one.
Count schema tokens as their own line item. Not "the system prompt." Serialise the tools array on its own and count it. That single number is the one that will change your mind, and it is the one nobody has to hand.
Budget in bytes, not in tool count. Sort your tools by serialised size and look at the top five. In most catalogues a handful of tools with long descriptions and deeply nested parameters account for a startling share of the bill, and tightening those descriptions is the cheapest win available.
Do the subtraction out loud. Context ceiling, minus reserved output, minus base prompt, equals what is left for the work. Write those four numbers down. If the remainder is smaller than a typical task's transcript, you have found your ceiling, and no amount of model upgrading will move it.
Scope tools per lane, not per system. Different jobs need different tools. A catalogue assembled for "everything this organisation does" is the wrong input to any single task. Per-lane profiles are straightforward to add and the blast radius is small.
Fail open, always. Every uncertainty in the curation path should resolve to the full tool surface. An optimisation that blocks work when it is confused is worse than the problem it solves.
Check the reliability side, not just the bill. If your tool-call error rate drops after curation, that was never a token optimisation. That was a correctness fix that happened to save money, and it is the more valuable of the two.
Know your one-step undo before you ship. A setting, not a deploy. You will want it at an inconvenient hour.
The manual on my wall is still 185 pages. It just is not read aloud before every sentence any more, and the person doing the job now gets handed the four pages that relate to the job.
The reason I find this one worth writing up is not the tokens saved. It is that I had written both numbers down myself, weeks apart, and never put them side by side. The measurement that mattered was free, it took an afternoon, and the only thing standing between me and it was that I had never thought to look at what I was actually sending.




