A method, not a platform
Build a yard you can break.
Then move the traffic off the old line.
A working copy of your core systems that holds no customer data. It is a practice yard — nothing in it is loaded, so a mistake costs a re-do. You use it to take the real line apart, one shipper at a time.
Your core systems are the old branch line. It is not hard to leave because of what it carries. It is hard to leave because every application still loads out on it, and nobody has a complete list of what else does. The batch job that has to finish before the other one. The field somebody reused for a second purpose in 2004. The retry behavior nobody designed. Every modernization program you have funded was, underneath, an attempt to buy back that missing list. One kind buys it with money. The other buys it with outages. There is a third way. Build a yard that carries nothing real, then work in it until the moves are routine.
Diagram: the old line drawn as a dense block of three hundred cells. Each cell is one connection point or one dependency — one more place something still loads out. Every cell is dark, because nothing has been catalogued yet. As the page goes on, cells light up as the traffic records explain them, change color as new sidings get built, and drop away once nothing loads on them.
Diagram, described in full: on the left, the protected side, holding Goliath as a block of three hundred cells plus six applications that still load out on him. In the centre, the Airlock, a boundary the customer runs, with a gate for things coming in, a holding area that returns rejected files along with the reason, and a single signed lane for the one thing that goes out. The key to the gate sits on the customer's side of the boundary. On the right, the clean build side: a store of waybills and traffic records, the Twin Compiler, a ledger of shippers not yet accounted for, and inside a fenced practice yard, the copy itself as a grid of tiles, each one marked with how realistic it is. Arrows run from the old line to the Airlock to the waybill store to the compiler, from the compiler to the ledger, from the compiler to the copy, and three loops come back: questions out to the people who can answer them, results from the yard returning as evidence, and live data returning from sidings already carrying traffic in production. As the page advances, Goliath's cells light up, change color, and fall away, while the copy's tiles assemble.
I
The old line
What it costs, why the two usual approaches lose to it, and the one thing worth doing first.
01The cold fact
The old line is expensive because nobody knows who loads on it
You are not paying for what the line carries. You are paying for the traffic nobody can fully account for.
The line items are knowable. MIPS, licence renewals, the annual vendor increase, support contracts, the headcount that reads the console. Every one of those is on a page somebody in this room signs. None of them is the reason it is still here.
What keeps it here is the batch job that has to finish before the other one and nobody remembers why. The copybook field reused for a second purpose in 2004. The retry behavior that was never designed, only observed. The reconciliation rule that lives in one analyst's spreadsheet. That knowledge is real. The company runs on it. And it sits with a handful of people, some of whom have already retired.
So every estimate you receive carries a premium for what the estimator could not see. There is no loss run for it. Severity is unknown, frequency is unknown, and the reserve is set by anecdote.
- The invoice: MIPS, licence, support, headcount
- Add the dependencies nobody documented
- Add the timing rules three people know
- The part nobody priced is now the biggest part
A line nobody has catalogued is an exposure nobody has reserved.
02The trap
Both of the usual approaches fail for the same reason
One spends the budget rebuilding track you already have. The other finds the shipper nobody listed during a month-end window.
The first approach rebuilds the whole thing before anything useful ships. Licensing, hardware, masked data, environments, a platform to hold all of it. The setup work becomes the project. The budget goes to proving you can stand up what you already have, and the real system keeps changing underneath the copy while the copy is being built.
The second approach goes straight at production. Dual-run, reconcile, cut over. It moves fast until the first undocumented dependency shows up in a window that cannot slip. Then the company learns the wrong lesson, which is that the old system is untouchable. The next attempt is smaller, slower, and harder to get approved.
The two look like opposites. They fail for the same reason. Both need you to know the system before you act on it, and neither one gives you a safe place to learn it.
- First approach: rebuild it all before anything ships
- The setup work becomes the project
- Second approach: dual-run, reconcile, cut over
- A hidden dependency shows up during month-end
- Both were paying for the same missing knowledge
One buys the missing knowledge with money. The other buys it with outages.
03The inversion
You do not need to know every shipper to start
What you needed first was never a complete map of the line. It was a yard where being wrong costs a re-do.
Both approaches treat a missing fact as a reason to stop. It is not. A missing fact is a piece of work. It has a name, a rank based on how much work is waiting behind it, and one person or one query that settles it. The unknowns become schedulable the moment you stop needing all of them up front.
What you need before any of that is somewhere to be wrong for free. Not a lower environment. Those carry production data, production passwords, and production consequences. A place with no member record in it, nothing downstream depending on it, and no reason to be careful when you break it. Shaped like the real thing, made of nothing that matters.
You are not mapping the whole line. You are finding the next shipper to move, moving it, and looking at what came off with it.
- Stop trying to document everything first
- Take one application's connections, not the whole map
- Unknowns become ranked work, not roadblocks
- Being wrong on this side costs nothing
Failure on purpose in the yard is failure you priced. Failure by accident on the line is failure that priced you.
II
The interchange
One gate, run by you, and the reason the argument over who holds the data never has to happen.
04The Airlock
You hold the key to the gate, not the vendor
The gate runs in your own environment, using the security tools you already require. That settles the question of who holds the data before anyone can raise it.
The boundary is a gate you operate, enforced by the security tools you already require. The vendor does not hold the key and cannot open the gate wider. That is why the conversation about custody never has to happen. It is settled by the way the thing is wired, not by contract language.
The rule is simple. The side that holds your data writes down what the system does, in enough detail to build from. The other side builds from that written record and never sees the original.
- The key stays on your side
- Your own security tools do the enforcing
- The vendor cannot open the gate wider
- Trust comes from the setup, not from a promise
Nothing in the yard is real. Everything that leaves it is.
05Three gates
Admitted, quarantined, or signed on the way out
Everything coming in gets screened. Anything turned away becomes a task, not a dead end. Only one thing goes out, and it is replacement code.
Coming in, a file is admitted only after a scan for passwords and keys, removal of PHI and production data, a licence and ownership check, a malware scan, and a policy check. Admitted: source code, tests, build definitions, schemas and DDL, copybooks, batch and job definitions, runbooks, tickets, documentation, interview transcripts, approved telemetry.
Never admitted: raw PHI, member records, production passwords, unscrubbed logs, software you are not licensed to move, and files whose origin nobody can name. No amount of schedule pressure changes that list.
Anything turned away is held with the reason attached, so the person who sent it can fix it and send it again instead of giving up. Going out there is exactly one lane, and only replacement code uses it: built the same way every time, tested, scanned, signed, traceable, approved.
- Scan for passwords, strip PHI, check licence and origin
- Admitted: code, schemas, copybooks, runbooks, interviews
- Refused: PHI, member records, passwords, files of unknown origin
- The reason is kept, so it can be fixed and sent again
- Going out: rebuildable, signed, traceable, approved
A file that gets turned away is not a dead end. It is a task with a reason attached.
III
The yard
Whatever you have goes in. A working model of the line and a ranked list of what is still missing come out.
06The Twin Compiler
What the build cannot explain becomes the work list
Drop in whatever you have. The build produces a working model of the line and a ranked list of what is still missing.
It takes anything. Application source, runbooks, incident tickets, DDL, copybooks, job schedules, EDI companion guides, a forty-minute interview with the engineer who owns the nightly cycle, a telemetry extract. Everything follows the same path. Facts you hand over. Facts pulled out of the files. Gaps filled in by software. One connected picture. A ledger of the pieces still missing. Then working copies of the connection points.
The model and the ledger are outputs, not documents. Run the build again and both get remade from the evidence, start to finish. Nobody fixes a wrong answer by editing the output. You add better evidence and rebuild. That is why an interview note counts as evidence, and why the thing does not have to be right the first time.
Every statement records how we know it — we have evidence for it, we worked it out, or we assumed it — and a guess never gets promoted into a fact. That is tracked statement by statement, not document by document, so one trustworthy runbook does not vouch for the paragraph somebody guessed at.
- Drop anything: copybooks, DDL, runbooks, tickets, interviews
- Pull out the facts, fill the gaps, link it into one model
- Mark every statement: evidence, worked out, or assumed
- Every rebuild remakes the model from the evidence
An estimate never books as a settled claim.
07The ledger
The ledger sets the switching order, and it asks the questions
Each missing piece is ranked by how much work is waiting on it, and comes with instructions for exactly one person.
There are three kinds, and the build tool names each one precisely. Something missing that stops the work: an unresolved symbol. Something we worked out that still needs a person to confirm it: a weak symbol. Something we are not confident about that got built anyway: a warning.
Each missing piece is ranked by how many connections and how much of the campaign are waiting on it. Each one comes with a request attached: a named person, told only what they need to know to answer it, or an automated lookup where that is acceptable. Nobody is handed a documentation assignment. They get one narrow question the build could not answer.
Disagreement is not a problem here. Two people telling you different things does not break the build. The disagreement lands on the ledger as a missing piece instead of one version quietly winning. No one person has to hold the whole picture, which is the only reason this still works after the people who knew it have gone.
- Missing, and it stops the work — unresolved symbol
- Worked out, still needs confirming — weak symbol
- Low confidence, built anyway — warning
- Ranked by how much work is waiting behind it
- Each line becomes one question for one person
Nobody is asked to document the line. They are asked the one question the build could not answer.
08The practice yard
Nothing in the yard is loaded, so nothing is off limits
The copy behaves like the real line, down to the retry behavior and the known bugs, and there is not one member record in it.
The build rebuilds everything the applications actually touch. ldb-compatible schemas and behavior, APIs, files and batch feeds, queues and event streams, stand-in sign-on, timing and ordering, retry and timeout behavior, capacity limits, known bugs, vendor quirks, and the downstream reconciliation rules that decide whether a cutover gets accepted.
It also rebuilds the company. The three-day approval, the handoff between two teams, the window that only opens on Sunday — all of it built into the copy as real steps, because a migration plan that leaves them out is a plan for a different company.
There is no customer data in it, no production password, and nothing in it that cannot be rebuilt, so we attack it with no restraint at all. Kill processes. Interrupt feeds. Run things out of order. Run it out of capacity. Cut a dependency nobody wrote down and watch what falls over. Replay an outage that really happened. Force a half-finished deployment. Test the rollback. See how far the damage spreads, then do it again.
- Schemas, feeds, queues, sign-on, bugs, vendor quirks
- Approvals and handoffs built in as real steps
- No data, no passwords, nothing that cannot be rebuilt
- Attack without restraint, then do it again
Every successful attack becomes a procedure. Every failed attack becomes evidence.
09Realism
Feed it more is a shopping list, not a slogan
Realism is graded one connection point at a time, not for the line as a whole, and each one names the exact evidence that would move it up a rung.
The ladder has five rungs, and every connection point sits on exactly one of them. Nobody claims a rung. Evidence earns it, and the evidence it would take is written right there. That is what lets you rank the intake queue. You know what a document is worth before anyone goes and gets it.
Anyone can read the progress from outside. Missing pieces going down, realism going up. Those are counts off a dashboard, not a green-yellow-red status a person picked while under pressure.
- It answers
- It answers correctly
- It behaves like the real thing, bugs and all
- It behaves on the real clock: windows, order, retries, timeouts
- It can replay a real outage on demand
The ledger is the loss run for unknowns you never had.
IV
One shipper at a time
One shipper at a time, rehearsed in the yard, until the last thing loading on Goliath is a line item someone can cancel.
10The campaigns
The right question is which shipper stops loading on it
Bring in one shipper, rebuild only the connections it touches, prove it matches in the yard, then move a single path in production.
A campaign takes one application. Bring its real code in through the gate. Rebuild only the connections that application actually touches, only as realistically as it needs. Run it against the copy and write down everything it does: queries, transactions, jobs, files, downstream effects. None of that is estimated. It is watched.
Then build the new siding, or the piece that sits in between, and prove in the yard that it behaves the same and returns the same answers. Put it in production alongside the old system and move exactly one thing. One read path, one write path, one workflow, one batch job. Watch what production tells you. Send the differences back through the gate. Repeat until that application does not need the old system for anything.
Each move takes away connections, transactions, batch jobs, extracts, operational know-how, vendor dependence, licence justification, and one more reason the old system has to stay the system of record. Far enough down the list, can we retire it stops being a strategy debate and becomes an inventory check.
- Bring one application in through the gate
- Rebuild only the connections it actually touches
- Write down its queries, jobs, files, downstream effects
- Prove it matches, then run it beside the old system
- Move one path, collect the differences, repeat
Goliath is not lifted. The traffic is moved off him, one shipper at a time.
11The loop
One loop sharpens the copy and empties the old line
Documents, results from the yard, and live production data all feed one process that makes the copy better and the old line lighter every time around.
It is one loop. Drop something in, run it through the gate, file it as evidence, rebuild, compare the new model and ledger against the last ones, send out questions for the new unknowns, get answers, rebuild. It works the same whether what you dropped in was a copybook or a forty-minute interview.
Three things feed it. Files admitted at the gate. Results from the practice yard, because an attack that fails teaches you something true about the real line. And data from every replacement already running in production, because something running live is the best evidence there is.
Every input moves you forward and never backward. It adds a fact, confirms a guess, kills a guess, or settles an unknown. That is what compounds. The copy gets more accurate at the same rate the old system becomes less necessary.
- Something in, rebuild, compare model and ledger
- Yard results come back as evidence
- Live production data comes back as evidence
- Disagreements land on the ledger, nothing wins quietly
- Copy sharpens, old system shrinks, same loop
Missing pieces down, realism up. That is the entire status report.
12The end state
The yard outlives the project that built it
None of this is a consulting deliverable. The yard, the ledger, and the playbooks keep working after the last invoice.
At the end, you own the copy, still running and still getting better. You own the ledger, the first complete account of what your core systems actually are, how each fact in it is known, and what is still unknown. You own the playbooks, every one of them proven in the yard before it touched production. And you own the replacement code, which left through your own gate under your own signature.
And the old system is shut off. Licence retired, disaster-recovery obligation retired, one less thing to audit, and the people who ran it moved to systems that do not need a specialist to interpret them. Your finance group can already put a number on that last part. The number was never what was missing. What was missing was a believable path to a date.
The last thing you keep is the team. The tooling does not allow the sloppy version of this work. Where a fact came from gets recorded when it is written down. Anything the build produced gets rebuilt, never hand-edited. Every operation is rehearsed before it is run. The rollback is written before the cutover. Your engineers pick up the discipline because there is no other way to run a campaign.
- The copy, still running, still improving
- The ledger: what is known, and what is not
- Playbooks proven in the yard before production
- Licence, recovery, and audit obligations retired
- A team that practiced on its own systems
Building it and learning it are the same job.
Goliath does not fall. He is emptied until nothing loads on him.
The traffic comes off him one shipper at a time, until the last thing loading on him is a line item someone can cancel. By then the decision to abandon the line is not a decision. It is a formality, signed by someone who is no longer nervous, and the rails come up.
Start with one shipper and one gate. Everything after that is repetition, and the repetition is the whole method.
- The yardPermanent, and still improving. It behaves like the line, holds no customer data, and it is yours after the project closes.
- The ledgerThe first honest account of who still ships on the line. Every fact says how we know it, and it is plain about what is still unknown.
- The playbooksProven before they were used. Every production operation rehearsed in the yard first.
- The codeShipped through your own gate. Built the same way every time, signed, traceable, approved.
- The retirementLicence, recovery, and audit obligations retired along with the old line.
- The teamPracticed on your own systems. Building it and learning it are the same job.