Skip to article
Blog

How to track Claude Code costs: /usage, the Console, ccusage, and Claudoscope

· Liran Baba · 8 min read

I built a cost tracker because one Tuesday cost $47 and nothing on my machine could tell me why. The tooling has improved since, some of it from Anthropic and some of it from people like me, and there are now four reasonable ways to answer "what is Claude Code costing me." They answer at different resolutions, and the right one depends on whether you want a number for this session, this month, this team, or this one runaway loop.

Everything below was checked against the Claude Code docs, the ccusage docs, and my own transcripts on September 29, 2026.

First, what you are paying for

If you use an API key, the Console, Bedrock, Vertex AI, or Foundry, you pay per token: input, output, cache writes, and cache reads, at different rates. Cache reads are the cheap one. On Sonnet a cached input token costs about a tenth of an uncached one ($0.30 versus $3.00 per million when I last checked my own numbers), which is why cache behavior shows up so hard in a bill. Extended thinking is billed as output tokens.

If you are on Pro, Max, Team, or Enterprise, you are not paying per token. Your seat has an allowance that resets on a rolling five-hour window and a weekly window, and Claude Code draws from it alongside chat and Cowork. Dollar figures still appear in /usage, but they are informational unless you turn on usage credits to keep working past the allowance, at which point spend becomes real again.

For a sense of scale, Anthropic's own docs put the enterprise average at about $13 per developer per active day and $150 to $250 per developer per month, with 90% of users under $30 a day. Yours will depend on model choice, codebase size, and how many sessions you leave open.

Option 1: /usage inside Claude Code

Type /usage in a session. The Session block at the top shows total cost, API and wall-clock duration, lines changed, and a per-model token breakdown:

Total cost:            $0.55
Total duration (API):  6m 20s
Total duration (wall): 6h 33m 10s
Usage by model:
   claude-sonnet-4-6:  1.2k input, 5.3k output, 940.0k cache read, 50.0k cache write ($0.55)

Claude Code computes that dollar figure locally from token counts at list price. If your organization has contracted rates, an admin can set the modelPricing managed setting and the line will say at your organization's configured rates. Since v2.1.211 the totals reset when you /clear; before that they accumulated for the life of the process, which confused a lot of people.

Since v2.1.251 there is also a Prompt cache (main) line: request count, share of input served from cache, number of misses with the likely cause, and whether the cache is currently warm. That single line explains most surprising bills. If it says 40% of your input is coming from cache, you are paying full price for the other 60% on every request.

Subscribers get an extra breakdown: usage attributed to skills, subagents, plugins, and individual MCP servers, plus flags for behaviors like long context or cache misses when one accounts for 10% or more of recent usage. The same cost figure is available in the status line.

What /usage cannot do: show you yesterday. It covers the current session on the current machine, and it is gone after /clear. There is no history, no per-project rollup, and nothing that fires when a session goes off the rails. For a pattern report rather than a token count, /insights writes an HTML analysis of your recent sessions to ~/.claude/usage-data/report.html.

Option 2: the Console and the admin dashboards

This is the only source that is actually your bill.

API and Console users get the usage page at platform.claude.com. Authenticating Claude Code with a Console account auto-creates a workspace called "Claude Code" so its spend is separated from your production API traffic, and you can set a spend limit on that workspace. The Console dashboard shows spend per member, and the Claude Code Analytics API returns the same daily per-user metrics for a script.

Team and Enterprise admins get a spend report in org analytics with CSV export, updated daily, plus the Enterprise Analytics API on the Enterprise plan. Bedrock, Vertex, and Foundry users get their cloud provider's billing console and, for anything per-user, OpenTelemetry export or a gateway.

What none of these give you is a session. The granularity is a user and a day. If you are a developer on an Enterprise API deployment, you may not have access to any of it, which is the exact situation that made me start writing a parser.

Option 3: ccusage

ccusage is a CLI that reads the transcripts Claude Code already writes to ~/.claude/projects/ and prints usage and estimated cost by day, week, month, session, or 5-hour block:

npx ccusage@latest daily
npx ccusage@latest session --json

It runs anywhere Node or Bun runs, emits JSON for scripting, covers 18 agent CLIs beyond Claude Code, and works offline with cached pricing. Its auto cost mode prefers the costUSD value Claude Code writes into some records and falls back to LiteLLM pricing tables when it is missing.

Two caveats from its own docs: costs are estimates, and tool API calls such as web search are not included. If you want the terminal number in a cron job or a dashboard, this is the tool. I wrote a longer comparison with Claudoscope, including when ccusage is the better choice.

Option 4: Claudoscope

Claudoscope is a native macOS menu bar app that watches the same directory as it changes. Cost-wise it gives you:

  • A live cost figure next to each running session in the menu bar.
  • Analytics by project, by model, and by day, with a cache hit-rate view and a what-if calculator that reprices Opus sessions at Sonnet rates.
  • Cost alerts on four rules: a single-session cap, rolling-window spend from 5 minutes to 4 hours, a daily total, and a monthly total. They arrive as a macOS notification plus a red menu bar dot, and re-fire at each doubling instead of continuously.
  • A read-only MCP server so you can ask Claude Code "what did I spend on this project last week, by model" and get an answer from your own local data.

The estimates are reconciled against real Anthropic and Vertex invoices. That work is where I learned that web search was not being billed at all: Claude Code records the count in toolUseResult.searchCount, and the documented server_tool_use.web_search_requests field is always zero in transcripts. It is a cent per search, which adds up on research-heavy days.

Because it reads files, it works on the Enterprise API where per-session spend is otherwise invisible. macOS 14 or later, Apple Silicon only, free, MIT licensed, no network access.

Why no estimate matches the invoice

All four of these are estimates except the Console, and even the Console's per-day view will not line up with a per-session tool to the cent. The usual reasons:

  • List price versus contract. Every local tool prices at list unless told otherwise.
  • Cache TTL. A one-hour cache write is priced at twice the input rate, a five-minute one at 1.25x. Subscriptions default to the one-hour lifetime; API keys default to five minutes.
  • Web search and other server tool fees, which most trackers skip.
  • Streaming partial records. The JSONL contains intermediate records with a null stop_reason. Sum them naively and you double-count; I shipped that bug and ran 1.5 to 2x over the invoice until I found it.
  • New models. A tracker's pricing table lags a model launch by days, and each project handles unknown ids differently.
  • Data residency. Responses billed at the 1.1x rate are multiplied in /usage since v2.1.239 but may not be in third-party tools.
  • Background requests. Conversation summaries for --resume and prompt suggestions cost a little, typically under $0.04 a session, and never appear in a transcript-based count.
  • Other devices. Local tools see this machine. The Console sees everything.

If two tools disagree by a few percent, that is expected. If they disagree by half, one of them has a bug, and it is worth finding out which.

Where the money actually goes

From my own transcripts, in rough order of surprise:

Short sessions cost more than long ones. Dozens of quick questions, each loading context cold and exiting before the cache paid for itself, cost me more than a three-hour refactor. Fifty of those beat the refactor easily.

Cache misses after compaction. When context shifts, the cache busts, and until I had a hit-rate chart I had no idea 30 to 40% of my input on some projects was uncached.

Instructions you pay for every message. Most CLAUDE.md files across my team were 2,000 to 5,000 tokens. That is context window on every request. The docs now suggest keeping it under 200 lines and moving workflow detail into skills that load on demand.

Sessions left open all day. Claude Code sends the full conversation on every request, so a one-line question in a session that has been open since morning still carries the whole morning. /clear between unrelated tasks, /rename first so you can find it again.

Subagents and agent teams. Each one has its own context window. Anthropic's docs put agent teams at roughly 7x a normal session when teammates run in plan mode.

A setup that has worked for me

Keep /usage open while you work, mostly for the prompt cache line. Turn on a monthly cost alert at whatever number would make you wince, and a rolling-window alert that catches a loop before it finishes. Check the Console once a month against the local estimate; if the gap is growing, something changed. And set cleanupPeriodDays higher than the default 30, because a cost number is only useful if the session behind it still exists when you go looking.

Install

Claudoscope is free, MIT licensed, macOS 14 or later, Apple Silicon:

brew tap cordwainersmith/claudoscope
brew install --cask claudoscope

Or take the DMG from the releases page. If ccusage is the better fit for you, it is one npx away and I use it too.

Start exploring your Claude Code sessions, free and open source.

Free, MIT-licensed, and maintained in the open. Requires macOS 14.0 (Sonoma) or later.