There is a lazy critique of agentic AI systems that goes: they're bloated, they burn context on nothing, the overhead is waste. Someone captured Claude Code's first API request off the wire — the actual JSON that leaves the client before a single token of your problem gets processed — and the temptation is to read the resulting numbers as an indictment. ~18,000 tokens, ~73KB, before the model has seen your repository, your question, or your intent. Three-quarters of that is tool definitions the model may never call.
That is the wrong reading. Sit with the data long enough and it inverts: the tax isn't waste, it's the price tag on the restraint you actually experience when you use the tool. What follows is Agent Adda's reverse-engineering pass on that captured request — and the point of view that falls out of it.
What actually got captured
The exercise intercepted a live Anthropic-compatible POST /v1/messages call — the literal payload Claude Code sends before doing anything.
Full JSON Schema + restraint prose. Dominates the floor.
Date reminder, your prompt, agent/skill catalog.
Harness, memory rules, environment facts.
Fig. 1 — Startup request anatomy (v2.1.260)
| Category | Items | Bytes | Est. tokens |
|---|---|---|---|
| Tool schemas | 25 | 55,262 | ~13,800 |
| Messages | — | 10,900 | ~2,700 |
| System blocks | 3 | 6,623 | ~1,650 |
| Meta | 5 | 216 | ~50 |
| Total | — | 73,396 | ~18,200 |
The first finding that matters: your project isn't in here. No repo contents, no bulk-loaded source, no README dump. Files load on demand, via Read / Glob / Grep, only when the model decides it needs them. What's sitting in that 18K-token floor is capability, not content — the standing inventory of things Claude could do, priced whether or not it does any of them.
Fig. 2 — Repo source is not in the floor; it arrives only via tool results
Where the weight sits: tools, not typing
Within tool definitions, prose descriptions outweigh machine-readable schema by roughly two to one — 62.6% description text against 33% schema JSON. That ratio concentrates in orchestration tools.
Workflow, ScheduleWakeup, SendMessage, Cron*, Agent, Task* — most of the tax.
Bash, Read, Edit, Write, NotebookEdit — cheap and high-frequency.
EnterWorktree / ExitWorktree — isolation, heavy on prose.
Fig. 3 — Tool-family weight in bytes (approx.)
| Rank | Tool | Bytes | Desc | Schema | % req |
|---|---|---|---|---|---|
| 1 | Workflow | 5,425 | 3,480 | 1,790 | 7.4% |
| 2 | ScheduleWakeup | 4,981 | 3,396 | 1,392 | 6.8% |
| 3 | SendMessage | 4,879 | 3,386 | 1,306 | 6.6% |
| 4 | CronCreate | 4,089 | 2,924 | 958 | 5.6% |
| 5 | EnterWorktree | 4,048 | 3,220 | 692 | 5.5% |
Five tools, roughly 20KB of prose between them — most of it not documenting a parameter. Core coding tools are not the expensive ones.
Tool descriptions aren't documentation
Here is the opening clause of the Workflow tool's description, in spirit: only call this when the user has explicitly opted in — a specific keyword, a direct ask in the user's own words, or a skill that tells you to. For any other task, even one that would clearly benefit from parallelism, do not call this tool.
"That is not a parameter spec. That is a restraint policy — shipped at the point of temptation."
ScheduleWakeup tells the model not to poll for work the harness will notify anyway, and to respect cache-window economics when sleeping. CronCreate tells it to avoid :00 and :30 because every other session making the same "reasonable" choice would thundering-herd the API — and to nudge off that mark even when the user asked for a round number.
Fig. 4 — Guardrail-as-affordance: constraint travels with the capability
The visible behavior — an agent that doesn't fork twelve subagents for a one-line question, that picks 8:57 instead of 9:00 for a digest — isn't emergent politeness. It's paid for, in tokens, on every turn. The elegance is a recurring cost, not a one-time training artifact.
What's in system + catalog
Tools dominate, but the remaining quarter is not empty. System blocks (~9%, ~1.6k tokens) break into a billing header, identity one-liner, and ~6.5KB interactive policy — where Memory + Harness alone are ~60% of that text. Messages (~15%) include your prompt plus an injected agent/skill catalog (~7.4KB / ~1.85k tokens) — sixteen named types the Agent / Skill tools can reach.
Project source, nested path-scoped rules (until those files are read), and full skill bodies (until a skill is invoked). The instinct that a big codebase is eating startup context is measurably wrong at the starting gun.
The corollary the report undersells
Cost hygiene levers are real: --bare, short CLAUDE.md, prune MCP, lean memory. Full vs bare is stark:
~73 KB. Real agent sessions with orchestration, web, tasks.
~13.5 KB. Bash + Read + Edit only. −82% startup.
Fig. 5 — Full vs --bare startup floor
But deferring full tool descriptions until first use isn't merely theoretical — it's shipping. Cron control, cross-session messaging, worktree management, most MCP connectors sit as name-only stubs until a discovery call pulls the full schema.
Fig. 6 — Pay the lean core every turn; pay orchestration when you reach for it
The real state of the art isn't "pay the full tax every time" — it's pay for a lean high-frequency core up front, and the heavier orchestration tax only in sessions that reach for it.
If you're building agent systems
Every tool you register is a standing tax on every call that agent's fleet makes. Five orchestration-shaped tools cost more than the entire core action layer combined. The design lesson is not "give agents fewer tools." It's separate the always-on core from the judgment-heavy periphery, and gate the periphery behind a discovery step.
Fig. 7 — Architecture for discipline that travels with capability
Builder checklist
Beautiful behavior is not free and it is not learned once. It is re-purchased, in tokens, every time the model looks at what it could do next. The report priced that purchase at roughly 18,000 tokens a session. What it's actually pricing is discipline.