DISPATCHES · Summit Cognitive

← All dispatches

ProvenanceJuly 27, 20265 min read

The model that cannot explain itself

We keep asking machines to tell us why they did what they did — as if the answer were hiding inside them, waiting to be drawn out. It is not hiding. In most cases it was never there.

There is a request we make of automated systems so often that we have stopped hearing how strange it is. A decision lands on someone — a loan declined, a claim flagged, a candidate filtered out — and we turn to the system and ask it to explain itself. Tell us why. And there is a whole industry now devoted to answering that question after the fact: tools that highlight which inputs mattered most, that produce a plausible-sounding narrative of the machine's reasoning, that reconstruct, from the outside, a story of how the output might have come about. We treat this as though we are interrogating a witness who knows the truth and can be made to give it up. But the witness, in this case, kept no memory of the event. What these tools produce is not a recovered reason. It is a reason invented on demand, after the outcome is already fixed, to fit an answer that was reached without one.

This matters because of a claim that sounds like wordplay but is not: when a system's output cannot be traced to a reason, the reason did not exist. Not "the reason existed but was lost." Not "the reason is in there somewhere and we lack the tools to see it." I mean the stronger thing. If nothing was recorded at the moment of decision that ties the outcome to the inputs, the rule, and the state of the world that produced it, then there was no reason in the sense that matters — no basis that can be examined, tested, or contested. There was a computation, and there was a result. A reason is not the computation. A reason is an account of the computation that can be held up and inspected, and an account that was never written is not an account you can go and find later.

The confusion runs deep because we borrow the word explain from the human case, where it means something different. When a person explains a decision, they are not fabricating a story to fit their action; they are — at least in the honest case — recalling a deliberation that actually occurred, with reasons they actually weighed. The explanation refers back to a real prior event. When a system generates an explanation for an output it produced without any such deliberation, the word is doing entirely different work. It is not recalling a reason. It is manufacturing one that is consistent with the result. And a manufactured reason that fits the output is exactly as available for a correct decision as for a wrong one, which is precisely why it can tell you nothing about which you are looking at.

An explanation generated after the fact is not the reason a decision was made. It is a story that happens to be compatible with the answer — and a story compatible with the answer is compatible with the wrong answer too.

Explainability is a property of the record

The move that dissolves the confusion is to stop treating explainability as a property of the model and start treating it as a property of the record. A model is explainable not because it can be induced to narrate itself, but because the system around it wrote down, at the time of decision, the things an explanation would need to refer to: which inputs were actually consulted, what version of the rule was in force, what the state of the world was when the decision was made, what the output was and on what basis. If those things were captured, an explanation is not a reconstruction — it is a reading of a record that already exists. If they were not captured, no amount of post-hoc cleverness recovers them, because they were never there to recover.

This reframes the whole problem in a useful way. The question to ask of an automated decision is not is the model interpretable — a property of the internal machinery, endlessly debatable and often beside the point — but did the system leave a record from which this particular decision can be traced to its reasons. The first question can be answered yes for a system that keeps nothing, and the yes buys you nothing when a real decision is challenged. The second question is answered decision by decision, in the affirmative only when the provenance was actually written down. A Decision Receipt is exactly this: not the model's story about itself, but the captured basis of a specific decision, preserved at the moment it was made, available to be examined by someone with standing to ask.

And this is why the after-the-fact explanation is not merely weaker than the record but actively worse. It offers the appearance of accountability while supplying none of its substance. A person shown a generated explanation feels they have been given a reason, and stops asking. But the explanation refers to nothing; it cannot be checked against what the system actually did, because what the system actually did was never preserved. The person has been handed a plausible narrative in place of evidence, and the plausibility is the trap — it discharges the demand for a reason without ever producing one.

What it costs to have no reason

Consider what follows when a decision genuinely has no traceable reason. It cannot be contested, because there is nothing specific to contest — only an output and a story that would shift to accommodate any objection. It cannot be corrected in general, because you cannot find the flaw in reasoning that was never recorded. It cannot even be defended honestly by the institution that made it, which is a fact institutions discover too late: when a regulator or a court asks why this decision came out this way, the generated explanation is not evidence of anything, and the honest answer — we don't know, we didn't keep it — is the answer of a system that was never accountable, only automated.

None of this requires the model to be simple, or interpretable, or slow. A system can use machinery no human could follow step by step and still leave a fully traceable record of each decision — because the record is not a transcript of the machinery, it is a capture of the things a reason must refer to. The demand is not that the model be legible on the inside. The demand is that the decision be traceable on the outside: that when it landed on someone, the system wrote down enough that the outcome can be tied back to its basis by a party who was not present and does not take the institution's word for it.

So the next time a system is asked to explain itself and obligingly produces an explanation, the right question is not whether the explanation sounds convincing. It is whether the explanation is a reading of something that was recorded at the time, or a story assembled afterward to fit a result. If it is the second, then what you are looking at is a model that cannot, in the only sense that matters, explain itself at all — and no eloquence in the after-the-fact account changes the fact that the reason it offers is one it did not have when it decided. Explainability was never a talent we could coax out of the machine. It was a discipline we had to build into the record, before the decision, or not at all.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.