Repository navigation
fix: price 1-hour TTL cache writes at 2x input (fixes #44) - #45
Merged
long-910 merged 1 commit intoJul 28, 2026
Merged
Conversation
All cache-creation tokens were charged at the 5-minute cache-write rate (cacheCreatePerMillion, 1.25x input). Claude Code writes most of its cache with a 1-hour TTL, which Anthropic bills at 2x input, so reported cost was low. calculateCost now reads the per-TTL breakdown from usage.cache_creation and prices each portion at its own rate. The 1-hour count is clamped to the aggregate so a malformed record cannot inflate the total, and entries with no breakdown keep the whole bucket on the 5-minute rate, which preserves the existing behaviour for older logs. The breakdown is also threaded through the three call sites that rebuild TokenUsage by hand (prediction, project cost, heatmap), which would otherwise have kept charging the 5-minute rate.
Owner
|
@MildlyMeticulous |
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #44.
What
calculateCostcharged every cache-creation token atcacheCreatePerMillion(3.75, i.e. 1.25x input). That rate is the 5-minute cache-write tier. Claude Code writes most of its cache with a 1-hour TTL, which Anthropic bills at 2x input, so reported cost came out low.The API already returns the split alongside the aggregate:
calculateCostnow reads it and prices each portion at its own rate.Scale
On a 425-file
~/.claude/projectscorpus, 168,726,465 of 216,447,883 cache-creation tokens (78%) were 1-hour TTL. Billing that share at 1.25x instead of 2x understates the cache-write component by 37.5%.Notes on the implementation
The 1-hour rate is derived as
inputPerMillion * 2rather than added as a fifth setting, so it follows a user'sinputPerMillionoverride and there's nothing new to configure. Happy to switch it to acacheCreate1hPerMillionconfig key instead if you'd prefer it explicit — that's a small change and I'll follow your call.Two edge cases from the issue are handled and covered by tests:
cache_creationobject keep the whole bucket on the 5-minute rate, so older logs behave exactly as beforeThe breakdown is also threaded through the three places that rebuild
TokenUsageby hand —prediction.ts,projectCost.tsandheatmap.ts. Without that, prediction, project cost and the heatmap would have kept charging the 5-minute rate while the status bar showed the corrected figure.projectCost.tsneeded itsrawUsagecast widened fromRecord<string, number>toPartial<TokenUsage>to carry the nested object.Checks
npm run lintandnpm testboth clean — 69 passing, including 5 new cases: the all-1h rate, a mixed bucket, the no-breakdown fallback, the clamp, and custom pricing.docs/DATA.mdandCHANGELOG.mdare updated to match.