A Date Picker Beat Seven Days of Parsing

I build a reporting tool that answers plain-English questions about a clinic's data. For seven days, September 22 to 28, getting it to read dates out of sentences was a standing line of work: "Last quarter." "In June." "The first half of the year."

On the 27th I asked for one behavior: stop asking "which July?" and answer at once. By eleven that night, five cold reviews had each found a new way for that behavior to print a wrong number, and each fix made it smaller. By morning it was firing on 5 of the 47 questions it was built for. Nobody decided that.

A Confident Wrong Number

On September 22, "last quarter" came back as the wrong period. I asked how sure the interpreter had been, and an agent replayed the question 17 times. All 17 started the period on the wrong date, and each gave that date 99.6 to 100 percent probability. The right start date got about one chance in a trillion.

My golden set couldn't see it. Every case in it pins absolute dates, so its 98.1 percent said nothing about this path.

The first fix was a question back: when the words are ambiguous, the tool asks which period and offers choices. That's safe, and it's slow. On the 27th a fresh held-out set scored 53.5 percent right on the first turn, meaning right before anyone answers a card. Of its 40 first-turn misses, 27 were the tool asking which period.

I decided that was too much asking. I told the session: "I would rather pop a data range up after the question than go down this path too far." I picked answering at once on the likeliest period and showing the range under the answer, with no "which July?" card.

Five Reviews

The first commit landed at 1:49 pm. Each round of the build ended with a cold review: a separate agent with no stake in the code, whose job was to make it print a wrong number.

Review one found four wrong-number paths and a false statement in the feature's own documentation. My answer, at about 2:40 pm: "keep it simple, just bare months and quarters."

Review two found four more classes. Around 4:25 a usage limit stopped the workflow mid-round, and the build session finished with a blanket rule as a fallback. The commit that landed it called the rule mine, and so did the list of known losses. I never said it. A later review's notes corrected the attribution.

Review three found a count asked of a rate being answered with the rate. Review four found a rate named by part of its name answering as a different rate. Review five found a status word in front of a record type counting the wrong records.

The must-fixes tapered: five, four, three, one, one. That looks like convergence, and it was, by subtraction. Every fix ended the same way, with the question going back to the old which-period card. The fifth landed at 10:42 pm, eight hours and fifty-three minutes after the first commit.

Overnight, with my "yeah," three Opus and two Sonnet agents widened the path again: half-years, "since" a day, two periods compared. Every guard from the day ran unmodified.

The Number Nobody Chose

The next morning, set v14, written blind for this feature, got its first run. It scored 39.5 to 40.7 percent on the first turn across two runs, against 33.7 before the feature. A day and a night had bought five or six answers out of 86, with zero wrong numbers on every run.

I told the session, "I don't think v14 should have gone so badly," and asked for an independent agent to check whether the build had drifted. One Opus agent, no context from the build, $4.83 of live spend. Verdict: partly drift, no regression.

Even if the interpreter had read every question perfectly, the at-once path answered 5 of the 47 questions that used to get a which-period card. A second agent traced the guards. Of 40 carded questions that could be judged, 37 would have been right if answered at once. It was the guards holding them back, not the readings. One guard read "by month" as a longer period.

The morning's first report had said the path fired on 0 of 100. It was 5: the field that report read marks only answers that ran on the wrong period.

Every round had passed the same gate: no new wrong numbers, nothing lost against main. A narrowing passes that gate by construction, because giving back a gain isn't a loss against main. The gate asked whether anything got worse. It never asked how much of the feature was left.

Move the Decision

At 7:30 that morning, while the drift review was running, I redirected the work: "I don't want to keep hammering this date issue when we can just force them to specify a data period." A reporting app should ask for a range, not parse "the first half of the year" out of a sentence.

The ask box now makes you pick a date range before the question is sent. It was in by 8:47.

On the same 86 questions, two runs: 80.2 to 81.4 percent right on the first turn, against 39.5 to 40.7. After one pick, 90.7 to 91.9 percent. Zero wrong in every column.

That isn't an accuracy comparison, and the drift review had said so before I had the number. The old path's score after one pick was 89.5 to 93.0 percent, the same band. The interpreter didn't get smarter. A pick in the old flow was always a pin, and the picker moved that step to the front. The harness also pins each question's correct range, so 81 percent is a ceiling. The honest name for it is the review's: right when the person pins the period.

Two checks on that number. Set v14 was written for the old feature, and 80 of its 86 questions restate the period in words, so a pinned range and the words never disagreed. An agent covered that by pinning the wrong preset, "last quarter," on all 100 questions: 9 right, none wrong, the rest scored neither. And the pin path drew its own cold reviews, five fix rounds by evening. Their findings turned into checks on a range that disagrees with the words rather than into a smaller feature, and on the runs re-measured, the first-turn score stayed between 80.2 and 82.6 percent.

What I Deleted

That evening my build dashboard asked whether to keep or retire the at-once path for the command line and the API. Every web question now carried a range, so the path served only those two doors. My answer: "retire it."

The commit removed it outright instead of flagging it off: 31 files, 422 lines added, 2,524 removed. The 101 plain regression cases that used to answer at once now take one pick when no range is pinned, and each is listed as a loss I chose, with my words as the reason. The next morning I made the command line and the API refuse a first question with no range: "require it."

What the Loop Was Missing

A free-text interpreter has to guess a fact the person already had. The cure was to take the guess out of the loop, not to guard it harder or hand it to a smarter reader.

A review loop built on "find me a wrong number" can only subtract. The gate beside it needs a second line: how much of the feature is left. Five of 47 was a number the frozen set could have given at any commit after 2:32 pm on the 27th. The wrong-number count was the one in the loop, and it stayed at zero.

-- Justin Higgins. Software Engineer, Midwest. Watched five reviews shrink a feature to 5 of 47, then replaced the parsing with a date picker.


Companion pieces: Everything Looks Like Everything - the refusals are the product, and this is the cost side of the checks that produce them. Comprehension Confidence Is Useless - a number that stayed healthy while it measured the wrong thing.

Reactions, disagreements, war stories: jchigg2000.dev@gmail.com