Cut Claude Code token usage by 61%

Reduce Claude Code token usage with context profiles, /compact and a narrower tool set: a measured guide that takes the first request from 34.4k to 13.3k.

Solviera Teknoloji 6 min read Türkçe oku

If you want to reduce Claude Code token usage, the prompt is rarely the place to start. The real weight is the baseline context added to every request before you type a single word. In long conversations it drives both cost and how fast you hit usage limits. This post explains where that baseline comes from, how we measured it, and which settings cut it by up to 61%.

Where do the tokens go?

When you message a coding agent, the model sees more than what you wrote. The request also contains:

  • The system prompt
  • Tool definitions (Bash, Read, Edit, Write, Glob, Grep and others)
  • User-level plugins, hooks and skills
  • Tool definitions from connected MCP servers
  • The project’s instruction files (CLAUDE.md, AGENTS.md)
  • The conversation so far

The first five are re-read on every turn, and the last one grows as you go. A small baseline adds up quickly over a hundred-turn session.

Measured: how much context is in the first request?

We measured the context sent to the model on the first request of a new session with Claude Code 2.1 and Codex 0.155:

ProfileClaude CodeCodex
Standard≈34.4k≈13.3k
Balanced≈23.7k (−31%)—
Lean≈13.3k (−61%)≈9.7k (−27%)

The balanced profile limits the tool set to Bash, Read, Edit, Write, Glob and Grep. The lean profile also skips user-level plugins, hooks, skills and external MCP servers. On Codex, the lean profile turns off apps and web search.

Settings that didn’t help

While measuring, we found that some flags did not shrink the first request on their own:

  • --strict-mcp-config
  • --exclude-dynamic-system-prompt-sections
  • --disable-slash-commands

Claude Code already loads those parts only when needed. The real savings come from narrowing the tool set and not loading user-level plugins, hooks and skills.

We deliberately left two options out: --safe-mode also turns off the project’s CLAUDE.md, and --bare only works with an API key. Both are too restrictive for daily use.

How to reduce Claude Code token usage, step by step

1. Give the task only the tools it needs

An agent updating docs or refactoring one file doesn’t need a browser, image tools or a dozen MCP servers. Start these jobs with the balanced or lean profile.

2. Attach MCP servers selectively

Every MCP server adds its tool definitions to the context. Give a server only to the agents that will use it, not to all of them. AgentVera’s MCP manager shows the approximate token cost of each tool when you test a connection, so you spot a heavy server before it bloats every turn.

3. Watch the context size

You can’t save what you can’t see. AgentVera shows each agent’s live context size in its pane, so you look at a number instead of guessing when a conversation got expensive.

4. Compact long conversations

/compact summarizes the conversation so far and shrinks every turn after it. Run it by hand, or set a threshold in AgentVera to compact automatically once context passes it.

5. Limit what moves between agents

When a flow passes one agent’s output to another, it doesn’t have to send all of it. The Transform node in flows can truncate a reply, keep the first or last N lines, or keep only the last code block, so the target agent’s context doesn’t fill with noise.

6. Polish prompts with a cheaper model

If you want to tidy a draft before sending it, you don’t need the expensive main model. AgentVera’s Revise runs claude -p --model sonnet with tools off and no session log. File contents are not sent; only the draft, the agent’s name and role, and the project name.

Choosing a profile

WorkSuggested profileWhy
Browser testing, MCP-heavy integrationStandardNeeds the full tool set
Everyday feature workBalancedCore file and shell tools are enough
Docs, small refactorsLeanSmallest baseline, cheapest turns
Long research sessionBalanced + /compact at a thresholdKeeps context growth in check

Check the numbers on your own setup

Your numbers depend on your plugins, skills and MCP servers. To see your own baseline, open a new session, send one short message and look at the context size of the first request. Then try the same with a narrower profile. The difference is what you pay on every turn of the session.

Working out the difference over a session

Percentages can feel abstract, so think in terms of a session. Say you run 40 turns with one agent. Leaving conversation history aside, the baseline is re-read on every turn.

ProfileBaseline on the first requestBaseline alone over 40 turns
Standard≈34.4k≈1.38M
Balanced≈23.7k≈948k
Lean≈13.3k≈532k

This table only shows the baseline; real usage is higher once conversation history, tool output and the model’s replies are added. Provider mechanisms such as caching can also affect billing. Still, the table makes one thing clear: a smaller baseline makes every turn cheaper, and the gap grows as the session gets longer.

Which profile for which job?

The profile isn’t a set-and-forget setting. It depends on the agent’s role:

  • Research agent: reads many files and may need a browser or docs MCPs. Start with the standard profile and automatic compaction at a threshold.
  • Implementation agent: mostly reads and writes files and runs shell commands. Balanced is enough.
  • Test or docs agent: works in a narrow area. Lean saves the most here.

With parallel agents these differences add up. Giving each agent a profile that fits its role, instead of standard for all three, makes the same limit last longer.

Setting automatic /compact at a threshold

Manual compaction is easy to forget. Automatic compaction is simple: once an agent’s context passes your threshold, /compact runs. When picking the threshold:

  • Too low: the agent is summarized often and may lose earlier details.
  • Too high: compaction comes late, and every turn until then is expensive.

In practice an earlier threshold works for long research sessions and a later one for short implementation sessions. Since AgentVera shows the live context size in the pane, you can tune the threshold by watching it.

Common mistakes

  • Every MCP for every agent. Unused tool definitions are read on every turn.
  • Never compacting. A session with hundreds of turns keeps carrying old details it no longer needs.
  • Passing output as is. Sending a long reply in full to another agent bloats that agent’s context too.
  • Only trimming the prompt. The prompt is usually a small part of the total.

Wrapping up

Most Claude Code token savings come from the baseline that’s re-read on every turn. Narrowing tools to the task, attaching MCP servers selectively and compacting at the right time lets you do the same work with less context. On the setup we measured, that meant 61% less on the first request.

AgentVera lets you set context profiles, the live context indicator and automatic /compact per agent. See token savings and the Claude Code page for details, or start on the free plan.

Questions

Why is Claude Code’s first request so large?

The first request carries the system prompt, tool definitions, user-level plugins, hooks, skills and MCP servers. In our measurements a standard setup came to about 34.4k tokens.

When should I use /compact?

When a conversation gets long and the earlier details are no longer needed. Compacting shrinks every following turn. AgentVera can run it by hand or automatically at a threshold you set.

Does the lean profile turn off the project’s CLAUDE.md?

No. The lean profile skips user-level plugins, hooks, skills and external MCP servers; the project’s CLAUDE.md is still read.

Are there savings for Codex too?

Yes. With Codex the lean profile took the first request from about 13.3k to 9.7k tokens, roughly 27% less.

Bring your agents to one desk.

Download AgentVera for free; your installed CLIs are ready to go.

More posts