The Meter That Refused to Guess
In August I published an essay about what two coding tools' cost meters didn't count. It ended on an admission: I'd read one tool's logs closely and hadn't yet done the same dig through my own Claude Code logs.
My own tools had been reading those logs the whole time. On October 1 three of them were fixed in one day, each because it had treated the transcript on disk as a ledger. This is the dig, done through the tools I built to read them. The code and the fixes were written with Claude, and the commit trailers say so.
A Conservative Default
claude-ace is a small public command-line tool, published on npm, that reads the session logs in ~/.claude/projects and prints a summary: token totals, top models, event types. In July an agent and I added a dollar estimate to the main branch. The commit says models the price table doesn't know "fall back to $0 and are flagged, never guessed." The README states the same rule as a feature, so that an unknown model never inflates the total. The repo's first tests included one for that fallback.
Never inflating the total sounds like the careful choice, but it's a choice about which way to be wrong.
A price table is a snapshot, and this one was dated June 24. By October it had no row for Opus 5.5, Sonnet 5.5, Fable 5.1, Mythos 5.1 or Opus 5. The commit that fixed the meter is co-signed by Opus 5.5, so the table it replaced couldn't price the model helping to write the fix.
The report did disclose it. An unpriced model shows as "(no price)" in its row, and a footnote counts how many unpriced models were treated as $0. The disclosure was accurate, but it sat under a grand total that looked finished. A zero is a guess too. It guesses the model cost nothing.
Four Fifths of the Files
The second mistake was the scan. The collector read one level: each project folder and the log files directly inside it. Claude Code writes subagent and workflow-agent transcripts in subdirectories under each session, and the tool never opened them.
The fix commit counts them: about 13.7 thousand of about 17.4 thousand log files on my machine, roughly four in five, sat in those subdirectories.
I have two more numbers for the same blind spot. A desk dashboard of mine that shows Claude Code activity counted 693 tool calls over 24 hours. Once it read the subagent transcripts it counted 6,059, so about nine in ten of the calls had been invisible.
And in my audit of a week of sessions, 1,647 of 1,899 log files were subagent transcripts. I set those aside on purpose, because that question was about sessions. Different windows and different units, same shape.
One Response, Several Lines
The third mistake is in the format. Claude Code writes one log line per content block of a response, and each line repeats that response's usage. A response with some text and two tool calls is three lines. The early lines carry a partial output count, and the last line carries the final one.
claude-ace summed usage per line, so every multi-block response was counted once per block. That error points up.
The cost meter in my ux-tournament skill, which is public in claude-code-dev, went the other way. It already walked the subagent directories and already deduplicated by message id. But it kept the first line it met for each id, and the first line is the wrong one. The example the fix leaves in a code comment is a first line reporting 5 output tokens and a last line reporting 552. Keeping the first means counting 5 where 552 happened. That error points down.
claude-ace now keeps the per-field maximum across the lines of one id, and the skill's meter keeps the line with the largest output count. Same log, same fact, two meters, opposite errors.
A fourth case runs the other way again. In June my Claude Code history dashboard, claude-code-log, ranked "longest missions" from subagent dispatches alone. Those cap around 25 minutes. The longest main-thread turns ran to about 69, so the list implied the longest thing that ever happened was a 25-minute task. The desk dashboard saw only the top level. This list saw only the nested files.
Wrong Both Ways
So the counter was wrong in both directions at once. Skipped files and zero-priced models pulled the total down. Summing every line pushed it up. Keeping the first line pulled output down. The old total wasn't an underestimate or an overestimate, just a number whose error I couldn't sign.
I'm not putting a before-and-after dollar figure in this essay. The README calls the dollar section "estimates from a built-in, point-in-time pricing table," and the rates in the code are standard API rates per million tokens. A total is what the tokens would cost at list price. It isn't a bill, and it isn't what I paid.
The table isn't trustworthy yet either. My roadmap says an agent sourced it and it was never checked against Anthropic's pricing page. The same commit changed one model's rates, and the open item says nobody has verified whether the old rates were wrong or the price moved.
One more bias still points down. On October 1, 98 percent of cache-creation tokens in a sample of the 150 newest session logs were the one-hour kind, and the report prices every cache write at the five-minute rate.
The Fix Is Not Published
The fixes are on main, but they aren't released. package.json says 0.2.0, the same version npm has, and the changelog lists the fixes under Unreleased.
The tagged 0.2.0 release from May has no cost section at all. It does have the one-level scan and the per-line sum, so its token totals carry the second and third mistakes, and neither is fixed there.
The unknown-model sentence is also still in the README. The fix refreshed the table and left the rule. The next model will ship after the table's date, and the rule will do the same thing again.
What I Would Check
Before you trust a total built from transcripts on disk, check three things. Count the files with a recursive walk and compare that to what the tool says it scanned. Pick one response, count its lines, and find which line holds the final number. Write down which way your default errs, and put that direction next to the number, not in a footnote.
-- Justin Higgins. Software Engineer, Midwest. Found his own cost meter wrong in both directions at once and has not yet published the fix.
Companion pieces: What the Meter Did Not Count - the audit this one is the sequel to. Three in Five Opus Runs Did Not Need Opus - the week of sessions where 1,647 of 1,899 files were set aside on purpose.
Reactions, disagreements, war stories: jchigg2000.dev@gmail.com