DISPATCHES · Summit Cognitive

← All dispatches

MethodThe LedgerJuly 27, 20266 min read

The credence good

Economists have a name for a service whose quality you cannot judge even after you have received it — a credence good, like a surgery or a legal defense — and a machine-made decision is exactly that, which explains both why the market fails and what fixes it.

When you buy an apple, you can inspect it before you pay: you turn it over, you check for bruises, you decide. When you buy a restaurant meal you cannot judge it in advance, but the first bite settles the question — you know, that evening, whether the money was well spent. Economists have tidy names for these. The apple is a search good, its quality assessable before purchase. The meal is an experience good, assessable after. Most of what we buy falls into one of these two bins, and the market for it works roughly the way the textbooks promise, because a buyer who can eventually tell good from bad will, over time, reward the good.

There is a third bin, and it is the uncomfortable one. Some things you cannot judge even after you have consumed them. The economists Darby and Karni named these credence goods in 1973, and the examples are the ones that make people uneasy for exactly the right reason: the car repair, the medical treatment, the legal defense, the financial advice. You paid, the service was rendered, you lived with the result — and you still cannot say whether you got what you needed. The mechanic replaced the part; was it failing, or was he selling you a part? The surgeon operated; would watchful waiting have served you as well? The outcome, good or bad, does not answer the question, because the outcome is compatible with too many stories about the decision behind it.

The good you cannot judge even after

A consequential machine-made decision belongs squarely in that third bin. A model denies your loan, flags your account, ranks your résumé to the bottom, prices your policy at a number you did not expect. Suppose the answer was wrong. What, exactly, tells you so? Not the outcome you live with — you were denied, and denial is consistent with a sound decision on a genuinely weak file, and equally consistent with a costly error on a strong one. The two are indistinguishable from where you stand. You cannot re-run your life with the loan granted and compare. And the party on the other side is in a strangely similar position: the buyer of the decision system, the institution that deployed it, often cannot tell either, because it sees the same outputs you do and has no independent read on whether each one was appropriate to its case.

This is what makes the decision a credence good rather than merely a hard-to-evaluate one. It is not that judging quality is expensive; it is that the raw material for the judgment — the fit between what was decided and what the case actually warranted — is not present in anything either party can observe. The result is legible. The reasoning that should justify it is not. You are asked to take the soundness of the decision on credence: on faith that the thing was done properly, because you have no means to check.

You cannot tell a good decision from a lucky one by its outcome, which is why a market for decisions, left to outcomes alone, cannot tell them apart either.

The failures credence goods breed

Credence-good markets misbehave in ways that are, by now, thoroughly documented, and they misbehave for structural reasons rather than moral ones. The first symptom is chronic mistrust: because the buyer cannot verify quality, every transaction carries a suspicion the seller cannot easily dispel, and honest sellers are tarred with the same brush as dishonest ones. The second is inefficiency in the service itself — the literature on expert markets describes both overtreatment, where the expert provides and charges for more than the case needs, and underservice, where the expert provides less than it needs and pockets the difference. The mechanic recommends the fuller job; the specialist skips the harder diagnosis. Neither the overserved nor the underserved customer can readily tell, which is precisely why both persist. The third, at the far end, is outright fraud: when misconduct cannot be detected even after the fact, the deterrent that detection would supply is simply absent.

The machine version reproduces all three. Mistrust, because affected people learn that a confident denial and a careless one arrive in identical envelopes. Overtreatment and underservice, because a system tuned to a metric can be systematically too aggressive or too lax in ways no individual outcome exposes. And room for fraud in the broad sense — decisions that serve the deployer's convenience or margin rather than the case, dressed in the same neutral output as decisions that were done right. The common root is that trust us, the outcome was fine cannot discipline any of this. It is not an answer; it is a request to stop asking. In a market where quality is unobservable, an assurance that quality was fine carries exactly as much information as its opposite would — which is to say, none. The seller who did well and the seller who did badly can both say it, with equal fluency, at equal cost.

The credence-good fix

The interesting part is that credence-good markets do not simply collapse. Societies have built escapes, and the escapes are instructive because the same ones keep appearing across very different trades. Reputation, so that a seller's history follows it and gives it something to lose. Licensing and certification, so that a credential vouches for competence the buyer cannot assess directly. Liability, so that a bad outcome can be traced to a responsible party and priced. And — the one that does the most work, and the one the others quietly depend on — verifiable records: the itemized invoice, the operative note, the case file, the documentation that lets an independent third party, after the fact, check whether the service was appropriate to the situation. The mechanic's market improves when a second garage can read the diagnostic log; medicine's improves when a chart lets a reviewer ask whether the treatment fit the presentation. What turns a credence good back toward an experience good is not the buyer's own eye, which cannot reach the relevant facts. It is a record that makes those facts checkable by someone whose job is to check.

An examinable decision record is that instrument for machine-made decisions. A Decision Receipt that carries the evidence actually consulted, the rules that were in force, and enough state to replay the decision is not a nicer explanation; it is the thing that converts an unverifiable output into one an independent party can audit for appropriateness after the fact. It gives the honest deployer a way to prove it did well — a way that the careless one cannot counterfeit, because the record either holds up to scrutiny or it does not. That asymmetry is the whole point: it is what a bare assurance lacks and a checkable record supplies.

Honesty requires one boundary, though, and it is a real one. Some decisions are made under irreducible uncertainty, where even a complete record cannot prove the choice was cosmically right — the future refused to cooperate, and no amount of documentation retrieves the counterfactual you were denied. A record does not pretend otherwise. What it establishes is narrower and more defensible: whether the decision was sound on what was known at the time — whether the evidence supported it, whether the rules were followed, whether a reasonable process produced it. That is not the same as being right, and it should not claim to be. But it is exactly the standard we already apply to the surgeon and the advocate, whom we judge not on whether the patient lived or the case was won, but on whether the decision was defensible given what was in front of them. The credence-good problem was never that experts must be infallible. It was that their work must be checkable. For the machine, as for the mechanic, that is what a record is for.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.