The threshold is the verdict
A model produces a score; a single cutoff line turns that score into a yes or a no. The cutoff is the real decision — and it is usually set invisibly, by someone other than the decider, and almost never recorded as the choice it is.
A scoring system rarely decides anything. It produces a number — a probability of default, a likelihood of fraud, a risk of recurrence — and a number, by itself, refuses nobody and approves nobody. Between the score and the outcome there is one more step, so small it barely registers as a step: a line is drawn, and everything above it is treated one way and everything below it another. That line is the cutoff, the threshold, the place where a continuous measure is forced into a binary verdict. We talk endlessly about the model that produced the score. We almost never talk about the line. And the line is where the decision actually happens.
This is easy to miss because the line looks like a technical setting rather than a judgment. It is a single value, set once, buried in a configuration file or a policy table, and it does not change from case to case the way a score does. It has the quiet, fixed quality of plumbing. But a threshold is not plumbing. It is the precise point at which an institution has decided how to trade one kind of error against another — how many people it is willing to wrongly refuse in order to avoid wrongly approving, or the reverse. Move the line a little in one direction and a class of people who were being approved is now refused. Move it the other way and the institution absorbs more of the risk it was trying to push onto applicants. The number does not decide that. A person decides that, by choosing where the line goes.
What makes this consequential rather than merely overlooked is that the choice is a moral one wearing technical clothes. Where you set the threshold determines who bears the cost of the system's inevitable errors. A model that is wrong some fraction of the time is wrong about someone, and the cutoff decides which someones — whether the errors land mostly on people the system wrongly flags or mostly on the institution that wrongly clears them. That is a distribution of harm. It is the kind of thing an institution should have to defend out loud. Instead it is set quietly, often by whoever was closest to the dashboard, and then it disappears into the machinery as if it had always been there.
Two people who are not the same person
It clarifies things to notice that the person who sets the threshold and the person who is accountable for the decision are usually not the same person, and frequently do not know each other exists. The threshold-setter is often technical — a data scientist tuning for a metric, an engineer translating a vague instruction like "keep fraud under control" into a specific operating point. The accountable party is the institution that will have to answer when a particular person is refused and asks why. The first chooses the line. The second is bound by it. And nothing in the ordinary flow forces those two to meet, which means the most consequential parameter in the whole system can be set by someone who will never have to defend it to the person it rules against.
This severance is the heart of the problem, and it mirrors a pattern these dispatches keep finding: a decision that no identifiable person experiences as a decision. The threshold-setter experiences it as optimization — nudging a value to make a chart look right. The accountable party experiences it as inheritance — a setting that was already there when they arrived. Neither experiences it as the act it is: a deliberate choice about whom to disadvantage. The decision falls into the gap between the two, and a decision in a gap is a decision no one has to own.
The model estimates. The threshold decides. We have been auditing the part that only estimates.
And the threshold tends to drift, which makes it worse. A line set once for a particular population keeps applying as the population changes, as the upstream model is retuned, as the meaning of a given score quietly shifts beneath it. A cutoff that was a defensible trade-off in one season becomes an indefensible one in the next, without anyone having touched it — the world moved under a fixed line. Because no one thinks of the line as a decision, no one revisits it as one. It is the unexamined verdict, reissued unchanged against people it was never calibrated for.
Make the line a recorded choice
The repair begins with a refusal to treat the threshold as a setting. It is a decision, and like any consequential decision it owes a record — not of the score it produced, but of itself. A complete account of an automated outcome should be able to answer three questions about the line that produced it. What was the threshold in force at the moment of this decision, stated as the choice it is and not buried as a constant? Who set it, and on what basis — what trade-off between which errors was it chosen to strike, and against what population? And when was it last examined against the world it is now being applied to? A system that cannot answer those questions has not recorded its decision. It has recorded only the arithmetic that led up to it.
None of this is exotic. The threshold is already a value the system holds; making it a recorded choice is a matter of capturing it as such, with its rationale and its date, alongside the score it acts on. The difficulty is not technical. It is that recording the line as a decision makes the trade-off visible, and a visible trade-off can be questioned. An institution that writes down "we set this cutoff here, accepting these refusals to avoid those approvals" has produced something an affected party can argue with. That is precisely the point, and precisely what the current invisibility spares the institution from. The line stays buried because buried lines cannot be contested.
So the next time a system tells someone no, the honest question is not only whether the model scored them correctly. It is whether anyone chose, and can defend, the line their score was measured against — whether the cutoff that condemned them was a deliberate, dated, ownable judgment or just a number someone left in a config file and forgot. The model is the part everyone watches. The threshold is the part that decides. An institution that can account for the first and not the second has audited the estimate and left the verdict in the dark.
— Dispatches · Summit Cognitive
Continue from here
Turn the argument into a practice.
Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.