Cost · 4 min read ·
Control AI coding costs and token use
Where AI coding tokens really go, how to measure usage per agent, six habits that cut it, and budgets that stop a runaway run before it drains your plan.
Where do the tokens actually go?
Mostly into context. Anthropic’s cost docs say token costs scale with context size: the more context Claude processes, the more tokens you use. Claude Code sends your full conversation with every request, and each time it uses a tool it sends another request carrying the results. A one-line question at the end of a long session still pays for the whole history.
On top of that come subagents, which send their own requests, MCP servers, whose tool names and instructions load into context, and thinking tokens, which Anthropic’s docs say are billed as output tokens.
OpenAI says the same about Codex in plainer words. Tasks that look similar can consume different amounts of your allowance, because model choice, context, reasoning, tool use, retrieval and caching all affect usage, so prompt length alone is not a reliable estimate.
How do you measure before you cut anything?
Look at the numbers your tools already report.
In Claude Code, /usage shows token statistics for the current session. On a paid plan it also shows plan usage bars and a breakdown of what share of recent usage came from skills, subagents, plugins and individual MCP servers, with flags when something like long context or cache misses accounts for 10% or more. /context shows what is taking up space. The dollar figure is computed locally from token counts at list price, so Anthropic calls it an estimate and sends you to the Console for authoritative billing.
In Codex, OpenAI’s docs point to a usage dashboard and /status in the CLI. Their advice is to check the dashboard every week or two to learn your pace.
In Gemini CLI, /stats shows token usage and, where caching applies, cached token savings. Per its token caching docs, caching is available with an API key or Vertex AI, not with a personal Google account sign-in.
Try this once: write down your usage before a typical task and after it. That tells you more than any rule of thumb.
Which habits cut usage the most?
These come from Anthropic’s cost guidance, and most apply to any agent.
One task per session. Clear the conversation between unrelated tasks, because stale context is paid for on every later request. In Claude Code that is /clear, with /rename first if you want to return.
Right-size the model. Anthropic says Sonnet handles most coding tasks well and costs less than Opus, which it suggests reserving for complex architectural decisions or multi-step reasoning. Lower the effort level for simple work.
Prune your tools. Disable MCP servers you are not using, and prefer plain CLI tools such as gh over an MCP server where one exists, because a CLI adds no per-tool listing to context.
Keep the instructions file short. Anthropic suggests keeping CLAUDE.md under 200 lines and moving workflow-specific instructions into skills, which load only when needed. Codex stops reading instruction files at 32 KiB by default.
Be specific. “Add input validation to the login function in auth.ts” makes the agent read less than “improve this codebase.” For changes across several files, plan before coding so a wrong direction costs a plan, not an implementation.
Stop early. When an agent heads the wrong way, interrupt it and rewind instead of letting it finish. Anthropic’s docs recommend this, and it is cheaper than a full wrong run.
How do you stop a runaway run?
A runaway run is usually a loop: the agent runs the same failing command again and again. Watch for the same error twice in a row, then stop it, rewind and rephrase.
For scripted runs, set hard caps. claude -p accepts --max-turns, which ends the run with an error when the limit is reached, and --max-budget-usd, which stops at a dollar amount. The budget uses Claude Code’s own estimate, so it can differ from your bill. Anthropic’s guide to building agents notes that it is common to include stopping conditions, such as a maximum number of iterations, to maintain control.
Remember that parallel work multiplies cost. Anthropic’s agent view docs say running ten agents in parallel uses quota roughly ten times as fast as running one, and its cost docs say agent teams use about seven times more tokens than standard sessions when teammates run in plan mode.
How do you set a budget you can trust?
On a subscription, your budget is a usage window, not dollars. Decide what share of the window each kind of task deserves, and look at the usage page at the end of each day for a week. You will find your real pace quickly.
On an API key, set the limit where the bill lives. Anthropic’s docs describe workspace spend limits in the Claude Console for total Claude Code spend. If you pay for usage beyond a subscription, Anthropic’s help center describes a monthly spend limit that you can set.
For single tasks, decide the cap before you start: this fix is worth at most this many turns, or this many minutes. A cap you chose in advance is easier to respect than one you invent while an agent is still running.
Budgets that live in the workspace
SwarmPane’s Token Guardian counts only usage your CLIs report, never estimates. Daily budgets stop runs before they overspend, with chat warned instead of blocked, and a run that loops on one failure is stopped with the evidence.
The Usage view shows one row per CLI from its own reports, labels estimates as estimates, and includes Codex’s live limits when Codex shares them.
SwarmPane runs the agent CLIs and accounts you already have, and your providers bill their own usage separately. Start with a 7-day trial for $1 and put a budget on your next run.