failure-signal · a story in three acts
privacy floor k ≥ 5
failure-signal / org diagnostics — decision-sheet edition

Your organization is already telling you where it hurts.

The failure-signal method reads the exhaust an organization produces anyway — counts, timing, team-to-team flow. Never content. Never individuals. This is the story of how it works, why it works, and exactly what happens when someone decides the floor is negotiable.

concept — simulated validation only team-level or coarser, k ≥ 5 it counts, it doesn't read
DocumentFS-STORY-01
RevisionA
DateAug 2026
StatusSimulated
SEC 01

What your organization emits

Every platform in the stack sheds operational exhaust as a side effect of normal work. The method collects it in three tiers, ordered by trust cost — and one rule holds at every tier.

TIER 1

Already in your stack

Jira · GitHub · Jenkins · ServiceNow · PagerDuty · Datadog · Snyk · Terraform

cycle timetime-to-first-reviewdeploy frequencyticket ping-pongvuln SLA agingapproval-gate dwell

NEVER: code contents, ticket text, vulnerability details beyond aging.

TIER 2

Flow metadata — where the collaboration graph lives

Slack · Teams · Exchange · Google Calendar · Zoom · Confluence · enterprise search

team↔team flow matrixresponse latencymeeting loadCC-list growthzero-result searches

NEVER: message text, subject lines, meeting content, document contents.

TIER 3

Context & people — aggregate only

Workday · SAP · Okta · Qualtrics / Culture Amp

attrition by teamtime-to-fillspan & layersleadership↔frontline deltaeNPS by team

NEVER: individual scores, named attrition, named survey responses. Badge and facilities data: excluded outright — highest sensitivity, least diagnostic gain.

team A → team B reply latency 41h team A → team B reply latency 41h

It counts. It doesn't read.

"counts, timing, latency, and team↔team flow only — no content, no individuals, k≥5 rolled up at collection time."

the one rule, verbatim — failure-signal-method.md, appendix A

SEC 02

Same shape, different organization

Why believe failure has recognizable signatures at all? Because history keeps drawing the same ones. Three documented failures, three industries, five decades — each mechanism settled by a public investigation, each mapping onto a telemetry shape it would have emitted in a modern stack. Would have — the method has never been run on any of these. That test is Phase 1, proposed and not yet built.

1986 · 2003 — AEROSPACE

NASA, twice

archetype: infoflow

O-ring erosion recurred flight after flight without a design fix; an engineering no-go was reversed under schedule pressure with no new data (Rogers Commission, 1986). Seventeen years later the same chain: imagery requests for a foam strike denied twice, closed as routine (CAIB, 2003).

silence

A warning that repeated — then went silent instead of escalating.

2011–2016 — BANKING

Wells Fargo

archetype: incentive

Millions of unauthorized accounts opened to hit cross-sell quotas; ~5,300 people fired for it over five years while the quota structure driving it stayed exactly in place (CFPB consent order, 2016).

compliance: red, for years production: green

Every dashboard green except the one nobody touched the incentive for.

1975–2012 — IMAGING

Kodak

archetype: frozenCore

A working digital camera built in-house in 1975; leadership shelved it to protect film margins. Patents piled up in the adjacent domain for two decades while investment at the core never moved. Chapter 11, 2012.

patents ↗ (dashed) core: flatline

The frozen system was Kodak's own future, sitting on its own patent shelf.

the shared shape: a true signal, and then nothing no structural change ever co-occurs no structural change ever co-occurs the shared shape: a true signal, and then nothing

In all three, the data existed long before the failure. The mechanism was never "no signal" — it was signal that never converted into structural change. That is what a detector can be shaped to look for.

SEC 03

The chain: localize → narrow → disambiguate

The detector is deliberately dumb — an anomaly scan, matched-filter scores against eight archetypes, and graph attribution. Every step is inspectable arithmetic. No model, no black box, nothing you couldn't audit with a spreadsheet.

SIMULATED STAGE 1

Localize

the cause (quiet)

Rank the most anomalous teams, divisions, and interfaces. Symptom location ≠ cause location: a dependency broker's victims stall loudly while the broker looks merely busy. Per-team scan finds the broker 6% of the time; the collaboration graph, 100%.

ILLUSTRATIVE STAGE 2

Narrow

broker
.92
silo
.34
overgov
.22
capability
.15
intake
.11
none of the above
.08

Score each anomaly against the eight-entry archetype library — evidence for and against every candidate, thresholds calibrated on healthy organizations at a 5% false-alarm rate. "None of the above" is always on the ballot; a library of eight must never force the nearest match.

SIMULATED STAGE 3

Disambiguate

One ambiguity survives every sensor: known-and-tolerated vs leaders-don't-know emit identical telemetry with opposite remedies. The split lives in what leadership believes — no system of record holds it. So: five frontline, two leaders, three questions. The perception delta is the sensor.

frontline 2.1
leadership 4.3

Delta ≥ 2 → information gap. Scores agree while metrics stay red → tolerated risk. If a follow-up interview runs past thirty minutes, the telemetry didn't do its job.

SEC 04

Why it works

Deliberately dumb, fully inspectable

Anomaly scan, matched filters, graph comparison. The people being diagnosed can audit every step — which is precisely what makes the diagnosis usable inside a real organization.

It knows how often it cries wolf

Thresholds come from healthy-organization calibration runs at a fixed 5% false-alarm rate. "How often is this wrong on a healthy org" is a stated number, not a vibe.

Cheapest sensor wins

Each failure class has a minimal sensor set that makes it visible; three of eight archetypes need only the delivery scorecard. And more data isn't free: adding the graph dropped single-team accuracy from 71% to 53%SIMULATED by adding competing hypotheses. Collect by failure class, not by appetite.

It prefers honest uncertainty

"Indistinguishable from X without Y" is a valid, preferred output, with the cheapest resolving signal named. The detector never presents a conclusion the data can't carry.

The wall is load-bearing

Teams under five roll up into their parent at derivation time, not report time — the small row never exists anywhere, so no audit, breach, or subpoena can ever produce it. Edge weights only, never content. No join path back to any individual record: the data model doesn't carry a key that could look. This isn't a compliance nicety bolted on top; it's what keeps collection politically survivable, sensors ungamed, and poll candor alive. Remove it and the instrument destroys itself — that's Act II.

ACT I

The bright path: structures get fixed

A declared integration dependency between two teams shows a flow score of 0.12SIMULATED — near-zero traffic on an edge that should be busy: two shared meetings a week where the baseline expects six, first-review latency of 41 hours against a portfolio median of 6. The detector's honest output: silo — or broker, indistinguishable without the third team's queue. The graph check comes back clean. Silo.

SIMULATED declared dependency — the edge that should be busy team A team B flow 0.12 → 0.87 team A team B flow 0.12 → 0.87 declared dependency — the edge that should be busy
diagnose the interface change the structure — joint queue, embed rotation re-measure the same edge edge recovers

Actions target structures, not people. Nobody was named. Nothing about any person was ever computed — not hidden, not access-controlled: never derived. The floor held, and the diagnosis still landed.

this is the entire pitch — the wall and the diagnosis are the same feature

ACT II

How it goes south: one favor at a time

Nobody attacks the wall head-on. It comes down as favors — each one small, each with a sponsor, each sounding perfectly reasonable in the meeting where it's asked. Watch the pill in the header.

Love the team view. Can we split it by squad? Some squads are only three people, but it's fine — we know them.

— quarterly business review, week 2

WHAT JUST HAPPENEDThe three-person row now exists. The floor was breached at derivation, which means it was breached everywhere — every export, every backup, every subpoena from now on. Trust doesn't degrade gracefully; it's a step function.

Who exactly is the bottleneck on that edge? Just names, just this once.

— escalation channel, week 5

team A ↔ team B · flow 0.12  →  D. Okafor ↔ P. Iyer · 14 unanswered threads
WHAT JUST HAPPENEDEdge weights grew names. The diagnostic just crossed from measuring an interface to surveilling two synthetic people — and "just this once" is now the precedent cited next time.

HR would like the per-engineer review-latency table for calibration season. It's data we already have, right?

— calibration prep, week 9

WHAT JUST HAPPENEDA count became a score. The number was built to describe a team's interface; it is now attached to individuals as evidence of performance, a purpose it was never validated for — with error bars nobody computed.

Set up the weekly export to the leadership folder. Legal signed off — it's just metadata.

— ops sync, week 12

WHAT JUST HAPPENEDA diagnostic became a dossier. Point-in-time inquiry became standing surveillance, and everyone in the org will figure that out faster than any comms plan can frame it.

Four favors, twelve weeks, zero villains. Every request came from someone trying to do their job well. That is how load-bearing walls actually come down.

the failure mode isn't malice — it's helpfulness with admin rights

ACT III

How it probably goes: it finds the bad apples

With the wall down, the same math that found silos now ranks people. It works immediately, it looks rigorous, and every row below is a misreading. All personas synthetic.

TALENT SIGNAL · WEEKLY EXPORT● LIVE
idsignal profileflagrecommendation
E-2114 low message volume · long uninterrupted focus blocks the deep worker. Quiet work reads as absence in metadata — by construction. LOW ENGAGEMENT manage out
E-0937 after-hours activity rising · meeting declines up 40% the one person covering an understaffed rotation — the finding was the rotation. FLIGHT RISK pre-emptive backfill
E-1408 dense cross-team edges · high reply centrality the mentor half the org quietly routes through. Coordination reads as empire-building. GATEKEEPER restrict scope
E-3021 review latency 3× peer median they review for five teams because they're the only one who can. The latency was a headcount finding. BOTTLENECK performance plan

These aren't glitches to be tuned out. Metadata scoring systematically misvalues whole work styles — deep work reads as low signal, coordination reads as centrality — and those proxy errors don't distribute evenly across roles, personalities, or demographics. The people flagged first are the people whose value was least legible to begin with.

QUARTER +1

The neutralizing

Quiet exits, restricted scopes, pre-emptive backfills. Nobody is told a score decided it. Leadership's dashboard has never looked cleaner, and attrition is officially "healthy turnover."

QUARTER +2

Goodhart arrives on schedule

Everyone now knows what the sensors reward. Message counts rise, real conversation moves off-platform, calendars fill with defensive meetings. The softest signals game first — polls, then flow, then delivery. The graph turns green and goes blind.

QUARTER +3

Candor dies where it mattered most

Survey deltas compress — not because leaders finally know, but because the frontline stopped saying. The one sensor that could split "known and tolerated" from "leaders have no idea" now reads calm in exactly the org where the gap is worst. The original silo, by the way, is still there.

It found its bad apples. Then it stopped finding anything true.

Run the detector on this organization now and it emits a familiar shape: leadership perception diverging from a frontline that's gone quiet, a green dashboard nobody believes. The tool that got pointed at people didn't just miss the real failure — it became one. Archetype: infoflow, with excellent production values.

EPILOGUE

Keep the wall

The floor is not the compliance section of the design. It is the design. Team-level or coarser, k ≥ 5 rolled up before anything is stored, edge weights only, no path back to a person — these are what keep collection survivable, sensors honest, and the poll's candor alive. Every capability in Act I exists because of the wall, and every failure in Act III follows from removing it.

And the individual-level signal? It's real — the method's own notes say so. The trust-compatible version flips it from surveillance to service: an opt-in mirror a person runs on their own exhaust, results belonging to them alone. The method's verdict, verbatim: "Keep the wall. Build the mirror first."

K ≥ 5 — THE WALL HOLDS · BUILD THE MIRROR FIRST