Cut token usage and cost in AI coding assistants without losing output quality.
Token usage is the bill — every turn re-sends your whole context window and you pay for it again.
Most of it is waste: filler prose, noisy logs, stale context, bloated instruction files.
Reach for this recipe when cost climbs faster than your output.
/context paints your context as a grid, so you cut the biggest consumers instead of guessing.
- Run
/contextin Claude Code. - Find the heavy blocks: tool schemas, instruction files, long file reads.
- Attack the biggest block first.
$ /context
MCP tool schemas ████████████ 28% ← biggest, cut first
file reads ████████ 19%
CLAUDE.md ████ 9%
(illustrative — replace with a screenshot of your real /context)
/cost tells you what a session actually costs and where the spend goes.
- Run
/cost(alias/usage). - Read the breakdown by skill, subagent, and MCP server.
- Re-run it after a change to confirm the spend really dropped.
$ /cost
Session: $0.42 · 1.2M tokens
By: subagents 38% · MCP 21% · main 41%
(illustrative — replace with a screenshot of your real /cost)
/insights analyses how you prompt — probably sub-optimal — so you fix the pattern, not one prompt. What you repeat every session belongs in the knowledge base, and the counter-intuitive habits you never noticed get surfaced so you can drop them.
- Run
/insights. - Move what you repeat into
CLAUDE.mdor a rule, and drop the habits it flags.
$ /insights
• You restate the test command in ~60% of sessions → put it in CLAUDE.md
• Long "summary" turns inflate output → ask for terse replies
(illustrative — replace with a screenshot of your real /insights)
Built-ins show one session; an analytics tool reads all your local logs and reveals where the bill truly sits. The lesson it surfaces: cache reads dwarf input + output, so caching, not generation, is most of the bill.
- Pick one:
prompt-analytics-for-claude-codeorccusage. - Run it —
uvx --from prompt-analytics-for-claude-code prompt-analytics summary— no setup, it parses~/.claude.
Your instruction file ships every turn, so each cut line saves on every message.
- Open
CLAUDE.md(or.github/copilot-instructions.md). - Cut it to essentials and add explicit conciseness rules.
- Model it on the instruction file the project-memory skill scaffolds, which is already written this way.
# CLAUDE.md — keep it terse
- Answer first. Lead with the result, then the reason. Drop pleasantries and hedging.
- No tool-call narration. No decorative tables or emoji unless they carry information.
- Keep verbatim: code, quoted errors, security warnings. Cut the rest.Approving the wrong direction burns tokens on rework, so let Claude explore read-only and propose a plan first.
- Press
Shift+Tabtwice to enter plan mode (or start withclaude --permission-mode plan). - Review the plan, then approve to switch to execution.
Shift+Tab Shift+Tab → ⏸ plan mode
Claude reads and proposes; no edits until you approve
See permission modes.
Stale early turns ride along and get re-billed every turn, so reset when the task changes.
- Finish a task, then run
/clearto drop the history and reload onlyCLAUDE.mdand memory. - Use
/compactinstead when you want to keep a summary of the same task.
$ /clear
history dropped → fresh window, CLAUDE.md + memory reloaded
See reduce token usage.
Compacting on your terms keeps what matters instead of letting auto-compaction guess.
- Watch context use and act around 60–70%.
- Run
/compactwith focus instructions naming what to keep.
$ /compact keep the repro steps and the failing test; drop the file dumps
Output is repetition you pay to generate, so cap the chatter. caveman forces short, filler-free replies (reported ~65% output cut, code intact) and auto-detects 30+ agents.
- Built-in route: set
"outputStyle": "concise"insettings.json. - Harder cut: install the
cavemanskill and invoke it like any skill —/caveman(or/caveman ultra); stop with "normal mode".
/caveman
before: The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.
after: New object ref each render. Inline object prop = new ref = re-render. Wrap in `useMemo`.
See output styles.
Test, install, and build logs flood context with lines the model never needs.
- Install a CLI proxy:
RTK(Rust) orSNIP(Go, YAML filters). - Prefix your command with it:
rtk cargo test.
flowchart LR
A["rtk cargo test"] --> R[RTK proxy]
R --> C["cargo test runs"]
C -->|"~25,000 tokens"| R
R -->|"filter · group · dedupe"| M["~2,500 tokens to model"]
Real saving: git push (15 lines, ~200 tokens) -> rtk git push (1 line, ~10 tokens).
Vendor dirs, build output, and secrets get pulled into context by accident. Deny reads on them so they stay out — they remain grep-able.
- In
settings.json, addRead(...)deny rules for large or sensitive paths. .claudeignoreis not shipped — deny rules are the official way.
{
"permissions": {
"deny": ["Read(./vendor/**)", "Read(./dist/**)", "Read(./.env)"]
}
}See permissions.
An MCP server's schema rides along every turn; a CLI costs tokens only when you call it. Newer MCP tooling adds tool/context selection that loads only the tools you pick — cheaper than before — but a CLI is still leaner and faster.
CLI (gh, acli, …) |
MCP server | |
|---|---|---|
| Token cost | a few, only when called | a schema every turn; less with tool/context selection |
| Speed | fastest | slower |
| Use when | a CLI exists | no CLI, or you need typed/live tool calls |
See mcp-installation.md.
You optimise what you can see, so expand the transcript to watch what each turn actually invokes and pulls into context.
- Press
Ctrl+Oto toggle the transcript — it shows detailed tool and skill usage and expands collapsed MCP calls. - Spot skills or tools that load on turns that don't need them, then scope or remove them.
Ctrl+O — transcript expanded
⎿ Skill: token-optimization
⎿ Called slack 3 times → expanded: 3 tool calls
(illustrative — replace with a screenshot of your real Ctrl+O transcript)
See the keyboard shortcuts.
The top model on routine work is wasted spend, so pin the model per skill or agent — cheap for routine, top-tier for hard reasoning.
- Set
modelin a skill's or an agent's frontmatter (haiku/sonnet/opus, a full id, orinherit). - Give routine scouts a small model; reserve
opusfor the hard reasoning.
# .claude/agents/explore.md — routine scouting on a cheap model
---
name: explore
description: Read-only codebase scout
tools: Read, Grep, Glob
model: haiku
---A skill's SKILL.md takes the same model: field (e.g. model: opus for a heavy step). See sub-agents and skills.
Test runs, log parsing, and wide exploration flood the main window. A subagent does it in its own context and hands back only a summary, so the bloat never lands in your session.
- Define an agent in
.claude/agents/<name>.mdwith only the tools it needs and a smallmodel. - Let it run the noisy op and return a short result.
# .claude/agents/test-runner.md
---
name: test-runner
description: Run the suite and return only the failures
tools: Bash, Read
model: haiku
---See sub-agents.
Cached input bills far cheaper, and cache reads are most of your tokens (step 4) — so don't throw the cache away mid-task. A model switch, an MCP connect or disconnect, or an effort change rebuilds it from scratch.
- Set the model and reasoning effort at the start of a task, not mid-work.
- Switch model or toggle MCP servers only at task boundaries, where a cache rebuild is acceptable.
mid-task model switch → cache invalidated → full re-bill
same model to a boundary → cache reads stay cheap
See prompt caching.
Extended reasoning can silently add thousands of tokens on tasks that don't need it.
- In
settings.json, setMAX_THINKING_TOKENSto0for routine work. - See the Claude Code settings docs.
{
"env": {
"MAX_THINKING_TOKENS": "0"
}
}Measure first, then stack the cheap wins — trimmed instructions, plan mode, a clean context, less chatter, filtered output — before the advanced routing, subagents, and cache discipline. Most of the bill is cache and repetition; cut those and the cost follows.
