Review the decisions, not the metrics
A dashboard tells you the aggregate is healthy and hides every individual injustice inside the average — govern by pulling real decisions and reading them, because the wrong ones do not show up in a number designed to summarize them away.
Open your last governance review and count how many actual decisions anyone read. In most oversight functions the honest number is zero. The meeting looked at approval rates, error rates, a fairness metric or two, a chart that was green last quarter and is green this one, and it adjourned satisfied. Nobody pulled a single real case — nobody read who was decided, on what evidence, under which rule, to what outcome. Oversight had become dashboard-watching, and the thing about a dashboard is that it will happily stay green while a knowable set of people are being handed indefensible verdicts underneath it. Stop reviewing the summary. Review the decisions.
The failure is subtle because the metrics are not lying. An approval rate of the expected shape, an error rate within tolerance, a disparity ratio inside the threshold — these can all be true at the same moment that a specific, describable group is being wronged in a way no one in the room can see. The reason is structural, not a matter of anyone being careless. An aggregate is a device for compression: it takes thousands of individual outcomes and returns one number, and the entire value of that number is that it has thrown the individual outcomes away. That is what it is for. But it means the edge case, the miscategorized applicant, the small population the model was never calibrated for — the exact decisions you most need to find — are the ones the metric is built to dissolve. They do not move the average enough to show. Summarizing is precisely how an individual injustice disappears.
So the directive is plain, and it is a governance directive, not an engineering one: build the reading of real decisions into oversight as a standing obligation, on a schedule, with the same seriousness you give the numbers. Treat "we only look at the dashboard" as a governance gap — a named deficiency in the review, not a defensible economy of attention. A function that has never once read a decision it is responsible for has not been overseeing the system. It has been watching a picture of the system, drawn by the system, in the resolution the system chose.
The green dashboard over the indefensible case
Consider what a metric can and cannot tell you. It can tell you that something has shifted at scale — that approvals fell across the board this month, that errors are climbing, that a gap between two groups has widened past where it sat before. That is real information and you want it. What the metric cannot tell you is whether any particular decision was right, because rightness lives in the specifics a metric has by design discarded: the evidence that was actually in front of the decision, the rule that was actually in force, whether the outcome actually followed from either. A number can be green while a decision is indefensible, and the two facts do not contradict each other, because they are answers to different questions. The metric answers how is the population doing on average. The case answers was this person decided correctly. Oversight that only ever asks the first has quietly stopped asking the second, and the second is the one an affected party, a regulator, or a court will ask.
An aggregate is built to make the individual case disappear, which is convenient for the metric and fatal for the person inside it.
This is why a systematically wrong decision can persist for a long time behind healthy numbers. If a subgroup is small, its bad outcomes are a rounding error in the aggregate — invisible on the chart, unmistakable in the file. The dashboard is not concealing the problem out of malice; it is doing exactly its job, which is to summarize, and summarizing a small injustice away is arithmetic. The reviewer who trusts the chart is not lazy. They are trusting an instrument to show them something it was never built to show. The correction is not a better metric. It is to go and look at what the metric cannot contain.
Read real decisions like case files
Looking means what it says. Pull a real sample of individual decision records on a schedule and read them the way a human reads a case — the way an appeals officer reads the file, or a reviewer reads a claim they might have to defend. Not a summary of the sample, not a chart derived from it: the records themselves, one at a time, each read for whether the decision it documents was actually sound given what was in front of it. A dozen decisions read closely will teach you things ten thousand decisions summarized cannot, because you are asking the question the average erased.
The sample has to be built with intent, because a purely random draw will over-represent the easy middle and under-represent exactly the decisions that go wrong. So weight it toward the hard cases on purpose: the edge cases near the threshold, the ones that were complained about or appealed, the outliers, the small populations the aggregate cannot see. Include enough ordinary decisions to stay honest about the baseline, but do not let the mundane center crowd out the margin, because the margin is where the indefensible verdicts live. The point of a case review is not to estimate a rate you already have on the dashboard. It is to find the decisions the rate is hiding.
None of this is possible unless the decision record exists and can be read. If all that survives a decision is its outcome and a score, there is nothing to review — you are back to counting outcomes, which is the metric again. The case review presupposes the work of instrumenting the decision: capturing the inputs it rested on, the policy in force, the moment a score became a verdict, and doing it in a form a person can actually sit down and read. A reason written for the person, not a debug log written for the machine, is what makes a case reviewable at all. Build the record readable, or the review has nothing to open.
Metrics and cases are complements
Be fair to the dashboard, because the argument is easy to overstate. Aggregates are not useless — they are necessary. They are how you detect drift, how you catch a scale problem the day it starts, how you notice that something moved before any single case could tell you. No case review reads fast enough or wide enough to see a population-level shift as it happens; only the metric does that. The claim is not that aggregates are worthless. It is that they are insufficient — that they answer one of the two questions oversight owes and are structurally blind to the other. Metrics find the problems that show up at scale. Case review finds the problems that hide inside the average. You need both, and they are not substitutes.
The trouble is that of the two halves, the case review is the one almost everyone skips. It is slower, it does not fit on a slide, it cannot be automated into a green light, and it asks a reviewer to sit with the discomfort of a specific person's specific outcome rather than the reassurance of a number in tolerance. So it gets deferred, and oversight quietly narrows to the half that is easy, and the function keeps meeting and keeps approving and never touches a decision. That narrowing is the gap. Name it as one. A governance regime that watches the dashboard and never reads the cases is not doing lighter oversight; it is doing oversight with a permanent blind spot exactly where the injustices are.
So carry one question into your next review, and refuse to let the charts answer it for you: when did we last read a real decision this system made, and did we deliberately read the hard ones? If the answer is that you have only ever read the metrics, you have not yet begun to govern. The number was designed to summarize the decisions away. Your job is to go and read the ones it summarized — before they become the case someone else reads back to you.
— Dispatches · Summit Cognitive
Continue from here
Turn the argument into a practice.
Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.