The Scorer Was Grading the Log Format
In May I wrote about the scorer I built for my Claude Code sessions. The design rule was that pattern checks own safety, a model judge owns quality, and a cap on the multiplier sits where nothing fluent can reach it. I still think that rule is right.
On September 13 I asked Claude for a self-paced loop to make the scorer more accurate and to re-check which signals still exist in current logs. The loop took forty of my recent sessions, re-ran the scorer on them after each change, and wrote the composite into a ledger. Sixteen of seventeen items shipped, as fourteen commits that day. The seventeenth is parked.
The rule held. The numbers were about the wrong things.
Every figure below comes from the pass with no model in it, the pattern checks alone, run on those same forty sessions.
Four Nouns, All Wrong
A scorer does arithmetic on sessions, events, prompts and tokens, so it has to know what those are. On the current logs, mine didn't.
Sessions. Each subagent's transcript is its own file, and the scorer counted every file as a session. It reported 275 sessions in forty. Folding the subagent files into their parents brought it to 40.
Events and tokens. One model response is written as several log lines, and each line repeats the same usage figures, so counting lines counted the same tokens again. Merging the split lines took total tokens from 173.9 million to 86.0 million. Events went from 91,056 to 26,630 once split lines were merged and the lines the harness writes for its own bookkeeping stopped counting. Parallel tool calls, which only show up when the calls sit in one response, went from a 0 percent rate to 10 percent.
Prompts. Task notifications, system reminders and the briefs one agent writes to another were all counted as things I typed. The scorer saw 889 human prompts averaging 3,254 characters. After the two fixes it saw 289, averaging 267.
Two Checks That Were Wrong in Opposite Directions
The permissions check looked for a command-line flag that never appears in the logs, which record permission mode as events instead. There were 1,324 bypass events in those logs, and the check saw none of them. The May essay says skipping permissions carries a penalty. It never fired.
The secret check failed the other way. It matched ordinary variable names in code my sessions wrote, 89 times. That zeroed Security Awareness, and the safety cap took the whole corpus down to 3x.
The cap did exactly what it was built to do, on a wrong input. Fixing the check moved Security Awareness from 0.00 to 3.00 and the band from 3x to 5x.
Model Selection had a third problem. Its top score, 5.00, came from the subagents. The session model is a fixed choice for me, so the tier mix it rewarded was theirs. It now scores how I dispatch work, and it dropped to 4.68.
The Headline Barely Moved
The composite went from 2.76 to 3.20. If that were the only number I looked at, I'd have concluded the recalibration changed little. The multiplier went from 3x to 8x.
The categories tell a different story. Seven of them are scored by the pattern checks, and the other four sit at a placeholder in this pass. Security Awareness rose by 3.00. Planning Rigor rose from 0.50 to 3.15. The signal it relied on had zero uses in the current logs, so it read near zero for everyone. Execution Order rose from 0.94 to 1.96. Prompt Engineering fell from 4.37 to 2.04 when agent prompts stopped counting as mine, then climbed to 3.20 after its own rewrite. Model Selection fell 0.32 and Code Safety fell from 5.00 to 4.72. Cost Efficiency ended at 4.50, exactly where it started, after passing through 5.00 and 3.50.
Three categories went up, three went down, and one ended flat, for a net 4.9 points across the seven. The composite is a mean over eleven categories, and four of those didn't move, so it shifted by 0.44. The errors pointed both ways, and the average hid them.
The multiplier is a band, so small moves in the composite can cross an edge. The first crossing was the secret check. One false positive had been setting the answer, and a cap can't tell a wrong input from a right one.
The Fix Was the Inputs
There's an obvious wrong fix: let something that can read the code look at those 89 hits and say they were variable names. That's a fluent instrument reaching into the deterministic one's territory, which is exactly what the May rule forbids. It would also fix one input and leave every other count where it was.
The structure didn't change. The eleven categories, the mean, the bands and the caps are the same as in May, and only what goes into them changed. The cap, the mean and the judge all sit downstream of the counts, and none of them can notice that a count is wrong. The judge changed in the same loop too, but that's a separate question.
What I Cannot Say Yet
The new numbers are better inputs, but they aren't validated. I haven't scored these sessions by hand to see whether I agree with 3.20. This is also one corpus of forty sessions, and the 275 is a fact about those forty and no others.
The tool exists to answer whether I'm getting better over time, and runs from before September 13 are on the old scale, so they can't be compared with runs after it.
The General Form
When a scorer is calibrated against a tool's log format, the format is part of the instrument. If the tool changes how it writes the logs, the instrument changes with it, and nothing reports an error.
Before trusting any score, count the nouns it's built on and compare each count to what the logs contain. Pin a fixed corpus, re-run it after every change, and write the number down. The ledger is the only reason I can say which fix moved what. And read the categories, not the headline, because a mean will hide two errors that point in opposite directions.
A cap that nothing fluent can reach is also a cap that nothing fluent can question. That's its job, and it means the inputs are the only place left to be right.
-- Justin Higgins. Software Engineer, Midwest. Recalibrated his own session scorer, kept the regex in charge, and fixed the inputs.
Companion pieces: The Judge Never Overrides the Regex - the May design this measures. Three Bugs in the Instrument - another instrument that reported wrong numbers as facts.
Reactions, disagreements, war stories: jchigg2000.dev@gmail.com