Survived Is Not Endorsed

Twice this year I pointed a swarm of agents at a pile of free government data and asked it to find something worth building. Both times the interesting output was not the findings. It was what the swarm said about its own process when I made it check.

The Skeptics

The second sweep ran 76 subagents in one sitting over public health-system data. Every number it produced got a label. Found, meaning published and cited with its date. Estimated, meaning arithmetic on found figures, with the arithmetic shown. Modeled, meaning it rests on assumptions, with the assumptions listed. Every candidate opportunity was attacked by three independent skeptics before it was sized: one on data reality, one on prior art, one on economics and regulation. The corrections they forced are stored next to the finding, not folded into it.

That sounds rigorous. The first version of the schema had one column for the skeptic verdict, and it ran two different things together.

Across the nine top picks, there was exactly one confirmed vote. Everything else that survived came through weakened. The skeptic forced corrections and the finding stood, in a smaller form. So the verdict column got split. Survived is not endorsed. A finding that made it through three attacks is a finding with three sets of scars, and the scars are the useful part. Storing it as a clean pass throws them away.

There is a related rule in the database README that I want to quote because I have never seen it in a research pipeline before. Mine the workflow journal for every agent result, including candidates the skeptics knocked down, whose verdicts never reach the workflow's return value. The losers are evidence too. A pipeline that only keeps winners cannot tell you why the winners won.

The Critic That Checked the Process

Two critics ran after the sweep. One checked rigor: are the numbers right. It pulled a landing page the sweep had cited and found it contained three percentages, none of which were the two the sweep had quoted. A load-bearing input had failed to check out three separate times and was being carried entirely on an upstream verifier's say-so. That is the ordinary kind of catch.

The other critic checked coverage, and it did not check any numbers at all. It checked whether the review process had been applied evenly. Of sixteen candidates, eleven got three skeptic lenses. One got one. Four got no verification votes at all. And one of those four survived into the final nine while the other three were cut. The cut line between them was drawn on something other than evidence.

I had not thought to ask that question. The numbers could all be right and the process could still be uneven. An uneven process means the ranking is partly a product of which candidates happened to get attacked. The coverage critic also had a structural read that stung: all nine survivors were cost-side, retrospective, and drawn from the same family of files. Not one touched revenue. Not one was forward-looking. It called the result a repricing shop's research agenda, not a health plan's, and it named the growth-side opportunity the sweep had walked past twice.

A critic that checks how evenly the rigor was spread, rather than the rigor itself. I do not have a name for that yet. It is the most useful agent I ran all year, and it cost less than lunch.

Forget Your Own Answer

The first sweep, in July, was the one that taught me to build the second one this way.

I asked a swarm to evaluate 66 free datasets and pick four a lean builder could turn into a business. It did. Then I ran the identical evaluation two days later with one change: ignore the final four and any rankings from before, start from the full dataset yourself.

It landed on a completely different final four. None of the first run's picks came back as top choices. Three orchestrated workflows, about 95 subagents, about 3.9 million tokens, 67 minutes of wall clock, and a different company at the end.

The second run also wrote down that it had failed its own stopping rule. The fourth seat was supposed to be stable across two consecutive rounds. It was not. Round one picked one candidate, round two a different one, round three the first one again. Every contender for that seat scored two out of five on seat-worthiness. The field had no confident fourth pick, and more rounds would thrash instead of settling, so it stopped and said so. The honest answer was that the fourth seat did not exist, and it shipped that as the answer instead of a fourth seat.

What the anchoring comparison showed, in one sentence from the second run: when you do not anchor on the prior answer, the strongest opportunities shift toward fast, self-serve, verified-demand report tools on separate buyers, and away from the slower, procurement-gated diligence plays the narrower run favored. The first run had been pulled toward its own earlier conclusion. I could not have seen that without running it blind.

The Uncomfortable Version

Everything above is about agents, and the reflex is to file it as an agent problem. I do not think it is.

Every research process I have watched inside a large company has the same shape. Some findings get attacked hard and survive weakened. Some never get attacked at all and survive clean. The clean ones look stronger. Nobody stores the losers. Nobody re-runs the analysis with the prior conclusion hidden, because the prior conclusion is in the deck the executive already saw.

The swarm did not invent any of that. It just ran fast enough, and logged thoroughly enough, that I could see it. Survived is not endorsed. The cut line has to be drawn on evidence. And if you want to know whether you anchored, you have to run it again without the anchor and be willing to get a different answer.

-- Justin Higgins. Software Engineer, Midwest. Ran the same evaluation twice, blind the second time, and got a different company.


Companion pieces: Twelve Days of Green Tests - what is the anchor I did not author? The Decision Log Ate the Project - three documents echoing one origin is not three confirmations.

Reactions, disagreements, war stories: jchigg2000.dev@gmail.com