DISPATCHES · Summit Cognitive

← All dispatches

MethodThe Long ReckoningJuly 27, 20266 min read

The controlled experiment

Medicine spent two centuries learning that a treatment which obviously works usually doesn't — and built the control group to protect itself from exactly the confident, plausible conclusion that machine decisions now reach at scale.

For most of its history, medicine was confident and wrong, and the two conditions were not unrelated. Physicians bled patients, blistered them, purged them, and dosed them with mercury, and they did these things not out of malice but out of evidence — the evidence of their own eyes. The patient who was bled often recovered. The physician saw the recovery follow the treatment and drew the natural conclusion. What no one could see, standing at a single bedside, was the patient who would have recovered anyway, or the one the bleeding killed a week later, or the hundred cases whose outcomes averaged out into a false impression of success. The clinician's confidence was real. It was also, for centuries, the single most reliable engine of harm in the profession.

The problem was never a shortage of intelligence or care. It was that human judgment is systematically fooled by a specific class of illusion. A patient at his worst tends to improve simply because he was measured at his worst — regression toward the mean, which flatters whatever was done at the low point. A patient who believes he is being treated often feels and even fares better — the placebo response, crediting the remedy for the belief. Sicker patients get one treatment and healthier ones another, so the treatments seem to differ when it was the patients who differed. And over all of it sits the plain human hunger to see a pattern, to have the cure work, to be the doctor who knew. Every one of these forces produces a result that looks exactly like a treatment working. None of them requires the treatment to work at all.

In 1747, aboard HMS Salisbury, the naval surgeon James Lind did something that reads as obvious now and was rare then: he took twelve sailors sick with scurvy, matched them as best he could, and gave six different remedies to six pairs at the same time, under the same diet and conditions. The pair given oranges and lemons recovered; the others, in the main, did not. What matters is not that Lind found the answer — he himself did not fully grasp what he had shown, and citrus took decades to become naval policy. What matters is the shape of the reasoning. He did not ask whether the sailors given citrus got better. He asked whether they got better than the ones who weren't. He built a comparison the citrus could have failed.

The remedy that obviously worked

That is the whole invention, and it took another two hundred years to mature. The comparison Lind sketched became the control group: a deliberately constructed set of cases treated the same in every respect but the one under test, so that the outcome you care about can be measured against the outcome you would have gotten anyway. Later came blinding — keeping patient and then physician ignorant of who received what, so that hope and expectation could not quietly do the work and take the credit. And in 1948 the British Medical Research Council published its trial of streptomycin for tuberculosis, now generally regarded as the first properly randomized controlled trial, in which chance alone decided who received the drug. Randomization was the final turn of the screw: it severed the last thread by which a hopeful clinician could, even unconsciously, sort the promising patients into the treatment arm.

Read as a technology, the controlled experiment is not really a way of finding cures. It is a discipline of institutionalized doubt. Its governing rule is a refusal: you do not believe the result that looks obviously right merely because it looks right. You assume you are being fooled — by the mean, by belief, by your own investment in the outcome — and you build, in advance, a comparison that would have exposed the fooling if it were happening. Then you believe only what survives that comparison. The control group is doubt made procedural, doubt you cannot skip because it is built into the structure of the test. Medicine's authority in the modern world rests less on any particular discovery than on this one refusal, repeated at scale for eighty years.

Medicine's great discovery was not a cure. It was the refusal to believe any cure that had not been given a fair chance to fail.

Confidence is the symptom, not the cure

Now consider what a machine decision system is. It is, at its core, a confident-result machine. It ingests inputs and emits an output — a score, a ranking, a determination, an answer — and it does so fluently, at volume, and very often with an attached expression of certainty. The output looks right. It is well-formed, internally consistent, responsive to the question. And the temptation it presents is precisely the temptation medicine spent two centuries learning to distrust: to accept the result because it is plausible and the system is sure. The confident physician at the bedside has been rebuilt in silicon and given a throughput the eighteenth century could not have imagined.

The decisive point is that confidence is not evidence. A system reporting that it is sure is telling you something about its internal state — its mood, if you like — and nothing about whether the world matches its output. Self-reported certainty is the clinician's hunch in new clothes: it feels like knowledge from the inside and is worth, on its own, exactly what the hunch was worth, which is to say nothing you could rely on. The failure mode that should frighten us is not the machine that is visibly unsure. It is the machine that is fluent, calibrated-sounding, and wrong — the confident-wrong result, produced a million times a day, each instance carrying the same air of having obviously worked. A result you have built no way to be wrong about is not a result you have tested. It is a hunch with a larger sample size of nobody checking.

A control group for a decision

The remedy is the one medicine already found, translated. A machine decision earns belief the way a treatment does: not by being fluent but by being subjected to a comparison it could have failed and surviving it. In practice this means building, deliberately and in advance, the thing the confident system would rather you skip — a way for the decision to be shown wrong. You hold out challenge cases and check whether the system's determinations track the ground truth or merely track its own confidence. You take a decision and replay it against a counterfactual: the same process, one input changed, to see whether the output moves the way it should or clings to its conclusion regardless. You ask, of every consequential output, the question the control group exists to force — what would have shown this wrong, and did we look?

This is why replaying a decision matters more than hearing it explained. An explanation is generated after the fact, under no obligation to match what actually drove the result; it is the system telling you, plausibly, why it was right — which is the clinician's confidence again, now articulate. A replay is the decision run a second time against what was actually known, a comparison that can diverge and thereby falsify. Reproducibility — the demand that a finding survive being produced again by someone determined to break it — is the standard medicine paid for in two centuries of graves, and it is the standard machine decisions are most tempted to skip because skipping it costs nothing until it costs everything. A Decision Receipt that carries enough of its own state to be replayed and contested is, in the exact sense Lind stumbled toward on the Salisbury, a decision that has been given a fair chance to fail. Believe the ones that survive. Distrust, on principle, the ones that were never at risk — however sure they sound.

— Dispatches · Summit Cognitive


Sources

  1. On James Lind's 1747 scurvy comparison aboard HMS Salisbury, its design and its slow adoption: "James Lind," Wikipedia; the James Lind Library, "Illustrating the development of fair tests of treatments."
  2. On the 1948 Medical Research Council streptomycin trial, widely regarded as the first randomized controlled trial: "Randomized controlled trial," Wikipedia; Medical Research Council, "Streptomycin treatment of pulmonary tuberculosis," British Medical Journal (1948).
  3. On the development of the control group, blinding, and placebo response: "Blinded experiment," Wikipedia; "Regression toward the mean," Wikipedia.

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.