DISPATCHES · Summit Cognitive

← All dispatches

MethodThe Field ManualJuly 27, 20265 min read

Record whether a human was really there

A human in the loop who clicks approve on ten thousand decisions a day is not oversight; they are a rubber stamp with a pulse — so record not just that a human was present but whether they actually engaged, or your accountability is a signature nobody read.

Somewhere in your system there is a field that records whether a human reviewed the decision, and it is almost certainly a boolean. Reviewed: true. A name, a timestamp, a checkbox flipped. When someone asks whether the automated call was overseen by a person, that field is the answer you will produce, and on its face it is a good answer — a human was in the loop, so the decision is a human decision, accountable in the ordinary way. But the field records the wrong thing. It records that a human was present. It says nothing about whether the human did anything, and those are not the same claim. The gap between them is where most of the oversight in automated decision-making actually lives, and it is a gap you can close, but only if you decide to instrument it.

“A human reviewed it” is the most abused sentence in automated decision-making. It is abused precisely because it is so useful: it converts a machine’s output into a human’s judgment, which shifts the accountability, satisfies the regulator, and shields the institution from the charge that a black box decided someone’s case. That is a great deal of work for four words to do, and the words do it whether or not any review occurred. So the directive is this: stop recording that a human was in the loop, and start recording whether the human was engaged. Capture the evidence that distinguishes a decision a person actually weighed from one they waved through. Otherwise the safeguard you are counting on is a signature, and you have no idea whether anyone read what they signed.

The rubber stamp with a pulse

Consider what nominal human oversight looks like when it fails, because it fails in a specific and recognizable shape. A reviewer sits at a queue. The machine has already scored each case and rendered a confident recommendation. The reviewer’s job is to approve or override. The queue is long, the throughput target is real, and the machine is right often enough that overriding it feels like second-guessing an expert. So the reviewer approves, and approves, and approves — hundreds of cases in an hour, thousands in a day, at a pace that makes genuine examination of any single one arithmetically impossible. On paper, every one of those decisions was reviewed by a human. In fact, none of them was.

Three forces converge here, and each is enough on its own to hollow out the review. The first is pace: a human cannot meaningfully evaluate a consequential decision in the seconds a high-volume queue allots, and no amount of good faith changes the math. The second is power: often the reviewer has no practical ability to say no — no time budgeted for dissent, no cover for the delay an override causes, no path that does not mark them as the obstacle. The third is automation bias, the well-documented human tendency to defer to a confident-looking machine, to treat its recommendation as the default and one’s own doubt as the thing that needs justifying. Put a tired person in front of a fluent system and ask them to disagree with it on their own initiative, at speed, with no support — and they will not. The review becomes a formality that the reviewer themselves could not tell you was a formality.

This is worse than having no human at all, and the reason is exact. An openly automated decision is at least honest about what it is; it invites the scrutiny that automation warrants. An unengaged human launders the machine’s decision as a considered human one. They add a signature and no scrutiny, and the signature is the problem — it is the thing that tells everyone downstream to stop looking. You have not added oversight. You have added a layer that defeats it while carrying its name.

A human who approves ten thousand decisions a day is not overseeing the machine; they are notarizing it, and a notary who reads nothing is just a stamp that costs a salary.

Record engagement, not just presence

The fix is not to demand more human review. It is to record enough about the review that happened to tell whether it was real, and then to let that record speak. Capture the metadata that separates examination from rubber-stamping. How long did the reviewer actually spend on this case before deciding — not clocked-in, but on this decision. What was actually put in front of them, and what did they open or expand versus leave collapsed and unread. Did they have, in front of them, the information they would need to disagree, and did they have the authority to act on the disagreement if they formed one. And across their queue, at what rate do they ever override the machine at all.

That last figure is the one that tells you the most, which is why it deserves its own instrument. The override rate — how often the human departs from the machine’s recommendation — is the clearest available signal of whether the human is a check or a conduit. A reviewer who overrides a meaningful fraction of the time is exercising independent judgment, whatever else is true. A reviewer whose approval rate sits at or near one hundred percent, sustained, at superhuman speed, is not overseeing anything. Treat that pattern as a red flag, not a success metric. It is easy to read a near-perfect concordance between human and machine as validation — look how well the reviewer agrees with the model. Read it the other way. Perfect agreement at impossible pace is the signature of a review that is not occurring, and it should trigger scrutiny of the oversight process itself, not a note of congratulation.

None of this requires naming the reviewer as the culprit; the point is not to punish the person at the end of an impossible queue. The point is to make the truth about the review legible — to the institution that relies on it, to the regulator that credits it, and to the person whose case it decided. A record that says a human was present is a claim. A record that says a human spent this long, saw this, and overrides at this rate is evidence, and the difference is the same difference that runs through everything in this series: a claim asks to be believed, and evidence can be checked.

Right-size it, and stay honest

Two honest qualifications, because the directive fails if you apply it without them. The first: not every decision needs deep human review, and fast approval is not always a failure. A high-volume, low-stakes flow where the cost of an error is small and reversible does not warrant a person weighing each case for minutes, and pretending it does would be its own kind of theater. The imperative is not that every decision deserve deliberation. It is that you stop claiming meaningful oversight where your own metadata shows there was none. Match the depth of real review to the stakes of the decision, and then let the record describe the review that actually corresponds to those stakes — no more, and no less. A reviewed:true on a life-altering denial and a reviewed:true on a trivial routing call are the same field telling two completely different lies, and the fix is to make the field carry the truth in both cases.

The second qualification cuts against the first, and you have to hold both. Measuring engagement can itself become surveillance theater. The moment time-on-task becomes a target, it becomes gameable — reviewers will learn to sit on a case for the required interval, to open the required panels, to produce the appearance of engagement that the metric rewards, and you will have instrumented a new performance instead of the underlying thing. So aim at honest indicators, not gameable ones. An override rate is harder to fake than a dwell time, because faking it means actually disagreeing with the machine, which is the behavior you wanted in the first place. The goal is not a dashboard that looks like oversight. It is a record that would let an outside party tell whether oversight occurred — which means the indicators have to be ones a reviewer under pressure could not cheaply counterfeit.

The whole discipline reduces to a single refusal. Refuse to let “a human reviewed it” stand as an unexamined claim in your system. If you assert human oversight as a safeguard, instrument it so the assertion can be checked, and be prepared for the check to fail — because sometimes it will, and the failing is the finding. The override rate is the tell. Engagement is the thing. Presence is not oversight, and a record that captures only presence is not documenting a safeguard. It is documenting a signature, and you owe the person on the other end of the decision the difference.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.