DISPATCHES · Summit Cognitive

← All dispatches

StandingJuly 27, 20265 min read

The clinician and the model

Average accuracy is a property of a population. Accountability is owed to a person.

A model arrives in a clinic with a number attached to it, and the number is almost always an average. It was right ninety-something percent of the time across some validation set; it outperformed the prior standard of care on a held-out cohort; it has a published figure and a confidence interval and, increasingly, a regulatory clearance that leans on those figures. All of this is real, and none of it is wrong. But the clinician standing in front of one patient is not treating a cohort. She is treating the person on the table, and the model's average tells her almost nothing about whether it is right here.

This is the gap I want to sit with, because it is where the quiet liability lives. We have built a generation of decision support that is accountable to a distribution and silent about the individual. The two are not the same kind of object. A distribution can be summarized, audited, and defended in aggregate. An individual is owed something different: a reason, in their own case, that they or their advocate can examine and, if it is wrong, contest. The model that scores well on the population can still hand the clinician a recommendation she has no way to inspect — and at that moment the average has done its job and abandoned her.

Accuracy on average is not a reason in particular

Consider what the clinician actually needs at the point of care. Not the model's lifetime batting average, which she cannot use to second-guess this output. What she needs is the basis for this recommendation: which features of this patient drove it, what evidence in the record it weighted, what it could not see, where it is operating near the edge of the data it was trained on. Those are the things a clinician integrates with everything else she knows — the patient in front of her, the history the model never read, the small contradicting sign that a number cannot carry. A recommendation she can see into becomes one more input to a human judgment. A recommendation she cannot see into becomes an instruction she must either obey or override blind.

And the override is where the liability quietly transfers. When a clinician departs from a model that turns out to have been right, she is exposed; when she follows a model that turns out to have been wrong, she is also exposed, because the law and the profession still locate the duty in her. She holds the responsibility for the decision either way. What the opaque model has done is hand her the responsibility while withholding the materials she would need to discharge it well. It has made her accountable for a recommendation whose reasoning she was never permitted to read. That is not decision support. That is risk transfer dressed as assistance.

A population can be defended by a statistic. A patient can only be defended by a reason they are allowed to see.

Standing belongs to the individual, not the cohort

There is a word for what the patient is owed here, and it is standing — the recognized position from which a person can ask a decision to account for itself. Standing is not satisfied by a good population-level metric, because the patient was never the population. The patient was one case, and the only account that answers to them is an account of that case: what the system relied on, what it weighed, what it would have taken to land somewhere else. A model that can produce a validation curve but cannot produce the basis for the decision it just made has excellent standing with a regulator and none at all with the person it affected.

This is why I keep returning to the idea that a consequential automated decision should leave behind a record built for the individual — the evidence actually consulted in this case, the rules active at the time, enough state that the decision could be replayed and, if it was wrong, shown to be wrong. Not a saliency picture generated after the fact to reassure, but the genuine basis, captured at the moment of decision and addressed to the party with the most at stake. A Decision Receipt of that kind is what converts a population-trained model into something a single patient can stand on. It moves the model from a thing that performs well in expectation to a thing that can answer a question in the particular.

None of this asks the model to be more accurate than it is. A model can be right far more often than any clinician and still owe its outputs a basis the clinician can read. The two virtues are independent: accuracy is about how often the answer is correct, and contestability is about whether the answer can be examined and challenged when it is not. We have spent a decade optimizing the first and treating the second as a courtesy, and the result is a class of systems that are statistically excellent and individually unaccountable — confident in aggregate, mute in the case that matters to the person living it.

What the clinic should require

So the requirement I would put on any model that informs a clinical decision is not a higher number on the validation set. It is that the model produce, for each recommendation, a basis the clinician can inspect now and the patient can contest later — what it relied on, what it could not see, and how far it was reaching beyond its evidence. That requirement does not slow good models down; it disciplines the ones that have been trading on their averages. A model that has a real reason for this case loses nothing by showing it. The only model that suffers under this rule is the one whose recommendation could not survive a careful clinician reading the reason behind it — and that is precisely the model a patient most needs the right to question.

The average will always be the easier thing to publish. It is clean, it is one number, and it lets a system claim accountability without ever facing the individual it answered for. But accountability was never owed to the cohort. It is owed to the patient on the table, and to the clinician who has to look at her and decide. Until the model can give those two people a reason they can hold, its accuracy is a property of a crowd that neither of them is standing in.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.