No single decision was wrong
A system can make a million individually defensible decisions and still do harm, because the harm lives in the pattern — and the pattern is the one thing nobody keeps a record of.
Imagine an automated decision system with a perfect paper trail. Every determination it makes carries a full record: the inputs as they stood, the version that fired, the threshold it applied, the authority behind it, enough state to replay. Pull any file at random and it holds. Contest any decision and the record survives the contest. By every standard I have argued for — admissibility, provenance, replay — this system is exemplary. Now let it run a million times. Somewhere in that million, without a single file going wrong, the system can be doing real harm: one cohort drifting toward worse outcomes, a feedback loop quietly narrowing who gets a yes, a threshold interacting with a population in a way nobody chose. Every decision defensible. The pattern wrong anyway.
This is not a paradox; it is a scale effect. The properties we demand of a decision — justified, examinable, replayable — are properties of instances. But some failures have no instance. A distribution can tilt without any individual determination being mistaken, the way a coastline can erode without any individual wave being remarkable. The wrong lives across the decisions, in their shape, and the shape appears in no record because the shape was never a thing anyone decided. Nobody chose the pattern. It emerged. And our entire machinery of accountability — the record, the challenge, the appeal, the replay — opens one decision at a time.
Consider who could catch it. The affected person holds exactly one case: theirs. From inside a single denial, a tilted distribution is invisible; their file is clean, their contest fails, correctly. The operator reviews escalations — a biased sample of the loud and the unlucky, not the population. The auditor samples cases and finds each one sound, because each one is sound. Everyone is looking through a magnifying glass at documents that are individually flawless, and the harm is only visible from the air.
A receipt per decision does not add up to a receipt for the pattern. The account of the many has to be built on purpose, or it does not exist.
Why the pattern has no record
It is tempting to think the aggregate account already exists implicitly — that since every decision is recorded, the pattern is just a query away. Sometimes, forensically, after the fact, that is even true: I have argued elsewhere that a thousand receipts read together can testify about the system that made them. But a queryable pile is not an account, for three reasons. First, there is no baseline: a measured distribution means nothing without a declared expectation to compare it against, and almost no one declares, at deployment, what distribution of outcomes they expect a model to produce. Whatever the pile shows later can be rationalized as normal, because nothing was ever on record as normal. Second, there is no cohort object: the population was never a recorded thing with an identity, a version history, and an owner — so the question "when did this pattern start?" has no anchor. Third, there is no pattern replay: you can rerun one decision against its preserved inputs, but answering "what changed?" for a distribution means rerunning the class under the prior configuration, and nothing in a per-decision record makes that possible.
So when the population-level question finally arrives — from a regulator, a litigant, an examiner, or your own risk team — the institution answers it with the tools of the case. It pulls individual records: clean. It replays a sample: confirmed. And it has answered nothing, because the question was never about any record that exists. Meeting a population-level question with case-level evidence is not a weak answer. It is a category error, performed under oath.
The aggregate account
What would it mean to keep a record of the many the way we keep a record of the one? Four things, none exotic. Declare the class: a decision made at volume is a decision class, a first-class object with a name, a version, and an owner — not an emergent property of infrastructure. Declare the expectation: at deployment, put on record the distribution of outcomes you expect this class to produce, over what population, measured how. This is the step institutions resist, because a declared expectation is a commitment — which is precisely why it is evidence. Record the actual: measure the distribution the class actually produces, on a cadence, in a form a stranger could read, retained like any other consequential record. Treat divergence as an event: when actual departs from expected, that is not a statistic, it is an occurrence — timestamped, owned, investigated, answered for, with the versioning to rerun the prior configuration and isolate what changed.
None of this replaces the per-decision record; the case and the class are different objects, and each needs its own account. And none of it makes a pattern defensible that isn't. What it does is make the pattern answerable — visible to its owner before it is visible to a plaintiff, explicable when the question comes, and correctable while correction is still cheap. The institutions that fare worst in pattern-level disputes are rarely the ones with the worst patterns. They are the ones who discovered their own aggregate behavior at the same time their examiner did, and had to reconstruct it, backward, under adversarial conditions, from a million clean files that answered every question except the one being asked.
My disclosure belongs here: I build in this space, and the argument obviously serves the thesis my company is built on. So test it against your own systems instead of taking my word. Pick your highest-volume automated decision. Ask what distribution of outcomes you expected it to produce this quarter — not roughly, but on record, from before the quarter started. If the answer is silence, then everything your system decides is being defended one case at a time, and the only account of the whole is the one someone else will eventually assemble for you. No single decision will have been wrong. That was never the question.
— Dispatches · Summit Cognitive
Continue from here
Turn the argument into a practice.
Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.