Point of View — Agent Adda

The tax on good behavior.

Why Claude Code costs ~18,000 tokens before it reads a single line of your code — and why that tax is restraint, not waste.

Capture Claude Code v2.1.260
Method Live POST /v1/messages
Floor ~73 KB · ~18,000 tokens
Date September 2026

There is a lazy critique of agentic AI systems that goes: they're bloated, they burn context on nothing, the overhead is waste. Someone captured Claude Code's first API request off the wire — the actual JSON that leaves the client before a single token of your problem gets processed — and the temptation is to read the resulting numbers as an indictment. ~18,000 tokens, ~73KB, before the model has seen your repository, your question, or your intent. Three-quarters of that is tool definitions the model may never call.

That is the wrong reading. Sit with the data long enough and it inverts: the tax isn't waste, it's the price tag on the restraint you actually experience when you use the tool. What follows is Agent Adda's reverse-engineering pass on that captured request — and the point of view that falls out of it.


The Capture

What actually got captured

The exercise intercepted a live Anthropic-compatible POST /v1/messages call — the literal payload Claude Code sends before doing anything.

75%
tools[] · 25 schemas

Full JSON Schema + restraint prose. Dominates the floor.

15%
messages[]

Date reminder, your prompt, agent/skill catalog.

9%
system[]

Harness, memory rules, environment facts.

Fig. 1 — Startup request anatomy (v2.1.260)

CategoryItemsBytesEst. tokens
Tool schemas2555,262~13,800
Messages—10,900~2,700
System blocks36,623~1,650
Meta5216~50
Total—73,396~18,200

The first finding that matters: your project isn't in here. No repo contents, no bulk-loaded source, no README dump. Files load on demand, via Read / Glob / Grep, only when the model decides it needs them. What's sitting in that 18K-token floor is capability, not content — the standing inventory of things Claude could do, priced whether or not it does any of them.

Fig. 2 — Repo source is not in the floor; it arrives only via tool results


Tool Tax

Where the weight sits: tools, not typing

Within tool definitions, prose descriptions outweigh machine-readable schema by roughly two to one — 62.6% description text against 33% schema JSON. That ratio concentrates in orchestration tools.

~35k
Orchestration / async

Workflow, ScheduleWakeup, SendMessage, Cron*, Agent, Task* — most of the tax.

~7.7k
Core coding

Bash, Read, Edit, Write, NotebookEdit — cheap and high-frequency.

~6.5k
Worktree

EnterWorktree / ExitWorktree — isolation, heavy on prose.

Fig. 3 — Tool-family weight in bytes (approx.)

RankToolBytesDescSchema% req
1Workflow5,4253,4801,7907.4%
2ScheduleWakeup4,9813,3961,3926.8%
3SendMessage4,8793,3861,3066.6%
4CronCreate4,0892,9249585.6%
5EnterWorktree4,0483,2206925.5%

Five tools, roughly 20KB of prose between them — most of it not documenting a parameter. Core coding tools are not the expensive ones.


Judgment Externalized

Tool descriptions aren't documentation

Here is the opening clause of the Workflow tool's description, in spirit: only call this when the user has explicitly opted in — a specific keyword, a direct ask in the user's own words, or a skill that tells you to. For any other task, even one that would clearly benefit from parallelism, do not call this tool.

"That is not a parameter spec. That is a restraint policy — shipped at the point of temptation."

ScheduleWakeup tells the model not to poll for work the harness will notify anyway, and to respect cache-window economics when sleeping. CronCreate tells it to avoid :00 and :30 because every other session making the same "reasonable" choice would thundering-herd the API — and to nudge off that mark even when the user asked for a round number.

Fig. 4 — Guardrail-as-affordance: constraint travels with the capability

The visible behavior — an agent that doesn't fork twelve subagents for a one-line question, that picks 8:57 instead of 9:00 for a digest — isn't emergent politeness. It's paid for, in tokens, on every turn. The elegance is a recurring cost, not a one-time training artifact.


The Other Quarter

What's in system + catalog

Tools dominate, but the remaining quarter is not empty. System blocks (~9%, ~1.6k tokens) break into a billing header, identity one-liner, and ~6.5KB interactive policy — where Memory + Harness alone are ~60% of that text. Messages (~15%) include your prompt plus an injected agent/skill catalog (~7.4KB / ~1.85k tokens) — sixteen named types the Agent / Skill tools can reach.

Still absent at t=0

Project source, nested path-scoped rules (until those files are read), and full skill bodies (until a skill is invoked). The instinct that a big codebase is eating startup context is measurably wrong at the starting gun.


Deferred Loading

The corollary the report undersells

Cost hygiene levers are real: --bare, short CLAUDE.md, prune MCP, lean memory. Full vs bare is stark:

~18k
Full · 25 tools

~73 KB. Real agent sessions with orchestration, web, tasks.

~3.2k
--bare · 3 tools

~13.5 KB. Bash + Read + Edit only. −82% startup.

Fig. 5 — Full vs --bare startup floor

But deferring full tool descriptions until first use isn't merely theoretical — it's shipping. Cron control, cross-session messaging, worktree management, most MCP connectors sit as name-only stubs until a discovery call pulls the full schema.

Fig. 6 — Pay the lean core every turn; pay orchestration when you reach for it

The real state of the art isn't "pay the full tax every time" — it's pay for a lean high-frequency core up front, and the heavier orchestration tax only in sessions that reach for it.


Agent Adda Read

If you're building agent systems

Every tool you register is a standing tax on every call that agent's fleet makes. Five orchestration-shaped tools cost more than the entire core action layer combined. The design lesson is not "give agents fewer tools." It's separate the always-on core from the judgment-heavy periphery, and gate the periphery behind a discovery step.

Fig. 7 — Architecture for discipline that travels with capability

Builder checklist

Do
Don't
Keep an always-on action core lean and resident
Register every orchestration primitive as always-on
Put spawn / schedule / message / fork behind discovery
Bury restraint only in a long system prompt
Encode judgment in the tool description at point of use
Assume training alone produces polite multi-agent behavior
Measure startup with a wire capture — know your floor
Assume a big repo is eating context at t=0
Treat bare-style modes as first-class for smoke tests
Optimize CLAUDE.md while ignoring MCP and skill catalogs

Beautiful behavior is not free and it is not learned once. It is re-purchased, in tokens, every time the model looks at what it could do next. The report priced that purchase at roughly 18,000 tokens a session. What it's actually pricing is discipline.