DISPATCHES · Summit Cognitive

← All dispatches

MethodJuly 27, 20265 min read

The error you are allowed to make

Every classifier makes two kinds of mistake, and a threshold decides which one is rarer. That choice is not arithmetic. It is a ruling about who pays for being wrong, written in a number so no one has to sign it.

Every system that sorts people into two piles can fail in two directions. It can flag someone who should have been left alone, or it can wave through someone it should have stopped. A false positive and a false negative. You cannot drive both to zero at once; push the threshold to catch more of the second and you produce more of the first, and the reverse. This is presented, almost always, as a tuning problem — a matter of finding the operating point that maximizes some aggregate score. It is dressed in the vocabulary of optimization, and the vocabulary does real work, because it makes the decision sound like something a curve decides rather than something a person chooses.

But the two errors do not land on the same people. A fraud filter tuned to miss nothing freezes the accounts of the innocent; tuned to inconvenience no one, it lets the theft through and the loss falls on whoever it falls on. A screening tool set to flag aggressively burdens the wrongly flagged; set to flag cautiously, it abandons the people the missed cases would have protected. The threshold is the seam between two populations. Moving it does not reduce harm. It relocates harm — from one group of human beings to another. The knob has a moral on the far end of it, and turning the knob is choosing who the moral falls on.

This is the part the optimization story is built to obscure. When you express the tradeoff as a single number to be maximized, you have already decided that one person's wrongful flag and another person's unprotected loss can be weighed on the same scale and summed. That summation is not given by the data. It is a value judgment smuggled in as a weighting term, and once it is inside the objective function it stops looking like a judgment at all. It looks like math. The person who set the weight has made a ruling about whose suffering counts for how much, and the ruling is now invisible, dissolved into a coefficient nobody will ever be asked to defend.

A threshold is a sentence handed down in advance, on a population that has not yet arrived, by someone who will never have to say it out loud.

None of this is an argument against thresholds. You cannot run a consequential system without one; refusing to choose an operating point is itself an operating point, usually a worse one. The argument is against the concealment. The allocation is unavoidable. Its anonymity is not. There is a difference between deciding that the wrongly flagged will bear the cost of catching more fraud and deciding it without anyone having to know that is what was decided. The first is governance. The second is governance with the accountability surgically removed, which is the condition most automated systems ship in.

Naming the population that pays

The test for whether an allocation has been made answerably is simple to state and uncomfortable to apply. Can you name the population that eats the error, and can someone be made to account for why it is them? Not in the abstract — not "users may occasionally be inconvenienced" — but concretely: at this operating point, these people, with these characteristics, absorb this kind of harm at this rate, and that was a choice, and here is who made it and on what grounds. A system that can produce that statement has converted a buried coefficient back into a decision. A system that cannot has hidden a ruling inside a threshold and thrown away the record of having ruled.

This is where decision admissibility stops being abstract. A consequential decision earns the right to count as real when its grounds can be reconstructed and contested by the person it landed on. A threshold decision is no different, except that it is made wholesale rather than retail — once, in advance, over everyone who will ever pass through. That makes the obligation larger, not smaller. The choice of operating point should carry its justification with it: who weighed the two errors against each other, what they took one to be worth relative to the other, and what would have to change about the world for the weighing to come out differently. Strip that away and you are left with a number that allocates suffering and answers to no one.

The asymmetry the number hides

The reason this matters now, and did not seem to before, is that the people on the receiving end have started to be able to see the seam. For most of the history of automated sorting, the population that ate the error could not even identify itself as a population. The wrongly flagged experienced their flag as bad luck. The unprotected experienced their loss as misfortune. Neither could point to the threshold and say: that line, set there, on purpose, is why this happened to me and not to someone else. The allocation was real, but it was illegible, and illegibility is the same thing as immunity.

That immunity is eroding. As decisions become reconstructable, the operating point becomes a visible artifact — a thing that was set, by someone, at a defensible or indefensible value. The question stops being whether a mistake occurred and becomes whether the distribution of mistakes was chosen honestly. An institution that can show its threshold was set deliberately, with the affected population named and the weighting owned, is in a fundamentally stronger position than one whose number simply fell out of an optimization run that maximized a quantity no one will admit to having selected. The first made a choice and can stand behind it. The second made the same choice and pretended it was arithmetic, and that pretense is exactly what will not survive contact with a single person who finally understands that the error they suffered was the one the system was tuned to let them keep.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.