The Test Logged the Miss and Went Green
In September I pulled a redaction library out of a gateway that sits in front of a small app and published it as a Go module called piidetect. It finds personal and health information in text, checks what it finds, and masks it. Version 0.1.0 went out on September 10 with a corpus of 26 marked-up test strings and a test that scored the library against them.
That test passed. Run in verbose mode, it also printed this line:
mrn_colon: missed MRN "MRN: 86753090"
The number is fake, but the miss was real. Feed v0.1.0 the sentence "See MRN: 86753090 for the prior imaging." and the sentence comes back unchanged.
The Floor That Allowed One
The test computed recall and precision across the corpus. It failed if recall fell below 0.95 or precision below 0.97, and it failed if a trap fired, meaning a string that looks sensitive and must not be flagged. A miss wasn't a failure. Misses went through a log call, not an error call. A comment in the file called the test the CI half of the recall ratchet.
The 26 entries held 20 values the regex recognizers were responsible for. One was a medical record number typed the way people type it: label, colon, space, digits. The recognizer had room for one separator character, and a colon followed by a space is two. So 19 of 20 were found, and 19 of 20 is exactly 0.95. The floor said fail below 0.95, and this wasn't below. Green.
A percentage sized to a corpus is a license for a number of misses, and at 20 values a floor of 0.95 licenses precisely one: one specific, named, known value, shipped for as long as nothing else broke.
The Loudest Form of Silence
Look at what the test did with the information. In a verbose run it printed the miss by name, with the value. In the CI workflow, which runs go test ./... -race -cover without the verbose flag, it printed ok. Go only shows the log lines of a passing test when asked, so on the runs that mattered, nobody had to ignore the miss. It was never shown.
The exit code was zero either way. The information was in the output and the exit code ignored it, and the exit code is what the pipeline reads.
For a classifier, 95 percent recall is a fine number. For redaction it isn't a rate of anything. Each miss is a value leaving unmasked, and masking values is the one thing the library exists to do.
What the Fix Changed
The fix is commit c393b5f, September 23, thirteen days after v0.1.0, committed with Claude as co-author and released as v0.1.1. Its message states the cause plainly: the test failed only on aggregate recall, and 19 of 20 is exactly the floor.
Three things changed in the test.
A miss fails on its own, with the value in the message. So does a claim the corpus didn't mark.
A hit has to cover the whole value. Before, any span that overlapped the value counted. That rule scores an email address like o'brien@example.com as found when the recognizer takes everything after the apostrophe. v0.1.0 returned o'[EMAIL], two characters of a name left in the output and counted as a hit. Under the new rule that's a miss.
The percentage floors stayed in the file, but the test no longer stands on them.
The Same Bug Was Everywhere
The commit message describes what came next as a sweep of every built-in recognizer for the same failure: a form the recognizer claims but silently misses. It found them. A toll-free number written with a leading 1. A card number followed by its expiry date. An IBAN printed in groups of four. A date of birth after its label in the ISO and day-first forms. An apostrophe in an email address.
None of them had a row in the corpus, and all of them are forms people actually write. Each fix came with corpus rows, and that one commit took the corpus from 26 entries to 44.
The Rule
The corpus is a regression list, not a sample. Every row is a value that must be masked, and every row fails by name.
Partial coverage counts as a miss. A value that's half masked has leaked the other half.
Percentages can sit on a dashboard next to the gate, but they can't be the gate. A threshold has a size, and a size is a count of misses you've agreed to ship.
Every miss you find becomes a row, so it can't come back unnoticed.
The Cheapest Audit
Search your own tests for a log call sitting near the word missed, skipped or known. Each hit is a failing test that was told to be quiet. It takes a minute.
In piidetect that search for Logf in the test files now returns nothing. The one skip left is the live integration test, which skips when the optional Presidio sidecar isn't running, and says so.
The README now states what the test enforces: every corpus value must be found and fully covered, and nothing unmarked may be claimed; one miss fails CI. That sentence is what keeps it fixed. It took thirteen days and one known value to write it down.
-- Justin Higgins. Software Engineer, Midwest. Gated a redactor on a percentage, and found the floor was sized for exactly one leak.
Companion pieces: Twelve Days of Green Tests - the oracle was circular there; here it was honest and the exit code ignored it. Spider-Man Is Not a Person - the false-positive side of redaction, with a green health check.
Reactions, disagreements, war stories: jchigg2000.dev@gmail.com