Skip to content

fix: price 1-hour TTL cache writes at 2x input (fixes #44) - #45

Merged
long-910 merged 1 commit into
long-910:mainfrom
MildlyMeticulous:fix/1h-cache-write-pricing
Jul 28, 2026
Merged

long-910 merged 1 commit into
long-910:mainfrom
MildlyMeticulous:fix/1h-cache-write-pricing

Conversation

@MildlyMeticulous

Copy link
Copy Markdown
Contributor

Fixes #44.

What

calculateCost charged every cache-creation token at cacheCreatePerMillion (3.75, i.e. 1.25x input). That rate is the 5-minute cache-write tier. Claude Code writes most of its cache with a 1-hour TTL, which Anthropic bills at 2x input, so reported cost came out low.

The API already returns the split alongside the aggregate:

"usage": {
  "cache_creation_input_tokens": 3904,
  "cache_creation": {
    "ephemeral_5m_input_tokens": 0,
    "ephemeral_1h_input_tokens": 3904
  }
}

calculateCost now reads it and prices each portion at its own rate.

Scale

On a 425-file ~/.claude/projects corpus, 168,726,465 of 216,447,883 cache-creation tokens (78%) were 1-hour TTL. Billing that share at 1.25x instead of 2x understates the cache-write component by 37.5%.

Notes on the implementation

The 1-hour rate is derived as inputPerMillion * 2 rather than added as a fifth setting, so it follows a user's inputPerMillion override and there's nothing new to configure. Happy to switch it to a cacheCreate1hPerMillion config key instead if you'd prefer it explicit — that's a small change and I'll follow your call.

Two edge cases from the issue are handled and covered by tests:

  • entries with no cache_creation object keep the whole bucket on the 5-minute rate, so older logs behave exactly as before
  • the 1-hour count is clamped to the aggregate, so a malformed record can't inflate the total (I saw 22 such records in ~41k)

The breakdown is also threaded through the three places that rebuild TokenUsage by hand — prediction.ts, projectCost.ts and heatmap.ts. Without that, prediction, project cost and the heatmap would have kept charging the 5-minute rate while the status bar showed the corrected figure. projectCost.ts needed its rawUsage cast widened from Record<string, number> to Partial<TokenUsage> to carry the nested object.

Checks

npm run lint and npm test both clean — 69 passing, including 5 new cases: the all-1h rate, a mixed bucket, the no-breakdown fallback, the clamp, and custom pricing.

docs/DATA.md and CHANGELOG.md are updated to match.

All cache-creation tokens were charged at the 5-minute cache-write rate
(cacheCreatePerMillion, 1.25x input). Claude Code writes most of its cache
with a 1-hour TTL, which Anthropic bills at 2x input, so reported cost was
low.

calculateCost now reads the per-TTL breakdown from usage.cache_creation and
prices each portion at its own rate. The 1-hour count is clamped to the
aggregate so a malformed record cannot inflate the total, and entries with no
breakdown keep the whole bucket on the 5-minute rate, which preserves the
existing behaviour for older logs.

The breakdown is also threaded through the three call sites that rebuild
TokenUsage by hand (prediction, project cost, heatmap), which would otherwise
have kept charging the 5-minute rate.
@long-910

Copy link
Copy Markdown
Owner

@MildlyMeticulous
Verified the fix — the 5m/1h split, the clamping, and the fallback for older logs all check out, and it's threaded through prediction/project-cost/heatmap consistently. Tests cover the right cases. Merging now — thanks for the clean fix! I'll get a release out with this shortly.

@long-910
long-910 merged commit fe1a6c3 into long-910:main Jul 28, 2026
@long-910 long-910 mentioned this pull request Jul 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cache-creation cost understated: 1-hour TTL cache writes billed at the 5-minute rate

2 participants