DISPATCHES · Summit Cognitive

← All dispatches

PilotsThe CaseJuly 27, 20268 min read

The shape of a first pilot

The first pilot is where the case for decision accountability either becomes real or becomes a slide. A good one is small, adversarial, and run on decisions that already happened — because the only proof that matters is whether a record can survive a hostile reading of a decision the institution already made.

There is a characteristic way that promising infrastructure dies in a first pilot, and it is not by failing. It is by succeeding at the wrong thing. The pilot is scoped to instrument everything, to touch many teams, to produce a broad dashboard that shows the capability is technically real — and it does show that, and everyone nods, and nothing is decided, because the pilot never posed a question whose answer would change anyone's mind. A decision-record layer is especially prone to this death, because it is easy to demonstrate that decisions can be captured and hard to resist the temptation to capture all of them at once. The result is a pilot that proves the least interesting claim — that recording is possible — while leaving the only claim that matters entirely untested. A good first pilot inverts this. It is deliberately narrow, deliberately adversarial, and built around a single question whose answer is a verdict, not a demo.

Pick one decision flow, and pick the one that scares you

The first design choice is scope, and the discipline is to resist breadth. A first pilot should cover exactly one consequential decision flow — one place where the institution makes a determination that affects someone, at some volume, on its own authority. Not the whole estate; one flow. Breadth is the enemy here because it dilutes the test: a pilot spread across ten low-stakes flows proves only that capture scales, which no one seriously doubted, while a pilot concentrated on one high-stakes flow proves whether the record actually holds where holding matters. The right flow to pick is not the easiest to instrument. It is the one whose decisions the institution would least want to be asked about — the flow where a challenge would be most expensive, most likely, or most embarrassing. That is where the value lives, and it is also where the pilot is most honest, because a record that survives the hard flow is credible everywhere and a record that only survives the easy flow has proven nothing about the exposure that motivated the exercise.

Define success as a single hostile question

The second choice, and the most important one, is the success criterion — and it should be a single question, phrased adversarially: take a real decision this flow already made, and defend it to someone who does not trust us. Not "can we capture the decision," but "can we produce an account that survives a skeptical, motivated outsider trying to poke holes in it." This reframing does almost all the work of a good pilot, because it forces the record to be tested against the standard it will actually face rather than the standard the institution finds comfortable. Everything the record must do — completeness, provenance, replayability, integrity, independence, readability — gets exercised automatically the moment you actually try to defend a specific decision to a specific skeptic, because the skeptic goes straight for whichever of those is weakest. The success criterion is not a metric. It is a confrontation, and it either resolves in the institution's favor or it does not.

A pilot that asks "can we record decisions?" always succeeds and never matters. A pilot that asks "can we defend one?" sometimes fails — which is exactly why it is worth running.

Run it on decisions that already happened

The third choice is temporal, and it is the one most pilots get backwards. The instinct is to stand up the record on new decisions going forward and wait to see what accumulates. But a forward-looking pilot cannot fail fast, because the interesting cases — the contested ones, the edge cases, the decisions someone actually wants to challenge — arrive on their own slow, unpredictable schedule, and the pilot ends before enough of them show up to test anything. The stronger design runs the record against decisions the institution already made — real historical cases, including the ones that were disputed, reversed, or uncomfortable. This does two things at once. It supplies, immediately, exactly the hard cases a forward pilot would have to wait months to encounter. And it directly tests the property that matters most and is hardest to fake: whether the record can faithfully reconstruct a decision after the fact, which is the entire job. A record that can only account for decisions made after it was installed has not been tested on the thing it exists to do. Reaching backward into real, messy, already-litigated cases is uncomfortable, which is the point — it is the only way to learn whether the record holds before an outsider makes you learn it the expensive way.

Fix the exit criteria before you start

The fourth choice is the exit, and it must be set in advance, because a pilot without a predefined verdict drifts into a permanent state of "promising." Before the pilot begins, name the small set of real decisions it will attempt to defend, name who will play the skeptic — ideally someone whose job is to find the hole, not to bless the tool — and name what counts as a pass and what counts as a fail. A pass is not "the vendor showed us a nice reconstruction." A pass is "we took these specific decisions, subjected each to a genuine adversarial reading, and the record held on all of them." A fail is any decision where the account could not be produced, could not be shown to reproduce the outcome, could not be shown to be unaltered, or could not be read by the outsider. Set that bar before you are emotionally invested in the result, because after you are invested, every partial success will look like enough, and partial success is exactly what an actual challenge will not accept.

Why this shape is the whole argument in miniature

The reason a first pilot deserves this much care is that it is the series' entire argument compressed into a single, testable exercise. The case for decision accountability rests on a claim the institution has never actually verified about itself: that if challenged, it could account for what its systems decided. Every essay in this sequence circles that claim — that it is urgent now, that the market for it is inevitable, that its absence is ruinously expensive, that it cannot credibly be built in-house, that a board will demand it, that a real trust layer must meet a hard specification. A well-scoped first pilot is where the institution stops arguing about the claim and finds out whether it is true, on its own decisions, before an outsider finds out for it. That is why the pilot should be narrow enough to finish, adversarial enough to fail, and pointed at decisions that already happened: not to show that recording is possible, but to answer, for one flow, on real cases, the only question that was ever at stake — can we stand behind what we decided? An institution that runs that pilot and passes has bought something real. An institution that runs it and fails has learned, cheaply and in private, what it would otherwise have learned expensively and in public. Both outcomes are worth more than a demo, which is why the first pilot is not a formality on the way to a purchase. It is the purchase decision, made honestly.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.