DISPATCHES · Summit Cognitive

← All dispatches

MethodThe Long ReckoningJuly 27, 20266 min read

The referee and the paper

Science does not trust a finding because a clever person announced it; it trusts it because independent referees who did not make the claim examined the method and tried to find it wanting first — scrutiny by disinterested strangers, built into the act of publishing.

Consider what has to happen before a scientific claim is allowed to become part of the record. A researcher, however brilliant, however certain, cannot simply publish a finding on the strength of being sure. The manuscript is sent to people the author did not choose and often does not know — experts in the same field, sometimes rivals — who are handed a single, adversarial assignment: read this carefully and tell us what is wrong with it. They probe the method. They ask whether the controls were adequate, whether the analysis supports the conclusion, whether an obvious alternative explanation was ruled out. Only after the work has survived this hostile reading is it admitted to the literature. The claim earns its place not by the confidence of the person making it but by outlasting the attempt of disinterested strangers to knock it down. That inversion — scrutiny before trust, performed by someone other than the author — is one of the more remarkable social technologies we have ever built, and we built it quietly, over centuries, mostly by learning what happens without it.

The lineage is older than the term. When the Royal Society began issuing its Philosophical Transactions in 1665 — arguably the first sustained scientific journal — the secretary who ran it did not simply print whatever was sent in. Submissions were read and assessed by the Society's members before they were communicated to the world; the institution interposed a layer of judgment between an author's claim and its publication. The practice matured slowly and unevenly. For a long time the editor's own discretion did much of the work, and formal external refereeing of every paper — sending it out to independent specialists as a routine condition of acceptance — did not become the standard, near-universal gate it is today until well into the twentieth century. But the principle was present at the creation: a finding does not enter the shared record on its author's say-so. It is examined first, by people whose job is to find the flaw.

What makes this examination valuable is not that the referees are smarter than the author. Often they are not. It is that they are structurally different from the author in three specific ways, and each difference does real work. They are independent — they did not make the claim, so they have no stake in its being true. Their charge is adversarial by design — the referee who finds a fatal flaw has done the job well, not badly, which means the incentive points toward scrutiny rather than applause. And their attention is trained on method, not merely on the result — the question is not "is this conclusion appealing?" but "was it reached in a way that could support it?" A pleasing result obtained by a broken method fails review. An inconvenient result obtained by a sound one passes. That is the whole discipline in one sentence.

A finding earns credence not from the authority of its author but from surviving the honest effort of a disinterested expert to prove it wrong.

Credence from surviving scrutiny, not from authority

It is worth being precise about what peer review actually certifies, because it is easy to over-claim. A refereed paper is not a true paper. Review does not verify that a finding is correct; it verifies that the finding has been examined by qualified people who did not produce it and who were looking for reasons to reject it, and that it survived. That is a weaker guarantee than certainty and a much stronger one than confidence. The value is entirely in the surviving. A claim that has passed through disinterested, method-focused, adversarial reading has been exposed to the most efficient error-finding mechanism we know — a motivated expert with no reason to be kind — and has not fallen. That is why the credential attaches to the claim rather than to the claimant. The most eminent scientist alive still submits to review; the unknown graduate student's paper, if the method holds, is admitted on the same terms. Standing is earned at the point of scrutiny, not carried in from reputation.

And here honesty compels the concession that keeps this from being a hymn. Peer review is imperfect, sometimes badly so. It misses fraud, because a determined fabricator can supply referees with data they have no way to independently reconstruct. It misses honest error that only later replication exposes. It can be conservative, punishing the genuinely novel result precisely because it is unfamiliar, and it can carry the biases of the people who happen to be doing the reviewing. None of this is a secret inside science; it is a standing complaint. But the complaint is not that review is worthless — it is that review is a check rather than a guarantee, and that we sometimes ask it to bear more weight than a check can carry. That distinction is the useful one. The institution's power was never that it made findings certain. It was that it made unexamined confidence insufficient — that it interposed, between a claim and its acceptance, a disinterested stranger obliged to look for the flaw. A flawed check that must be passed is worth more than a perfect standard that is merely asserted.

Refereeing the decision system

Now turn the lens, and notice how little of this survives the move from the scientist to the machine. A consequential decision system — one that scores a person, sorts a claim, flags a transaction, gates an opportunity — is typically admitted to action with nothing resembling refereeing. No independent expert who did not build it is handed the method by which a whole class of these decisions is made and tasked with finding what is wrong with it before the system is trusted at scale. There is no reviewer whose job is to reject it. We publish the decision straight to the world on the author's confidence — the confidence of the team that built the model, measured against benchmarks the team also chose. The claim goes to the record unrefereed, and the record here is not a journal that scholars will argue with over years; it is a set of decisions already landing on people's lives.

Peer review is the model for what is missing, and the analogy is exact in the parts that matter. What a consequential decision system owes the world, before and while it is relied upon, is independent, method-focused, adversarial pre-trust scrutiny: examination of how the system decides — not a demo of a few outputs it is proud of — conducted by disinterested experts who did not build it and who are specifically charged with finding it wanting. Independent, because the builder's own evaluation has the author's stake in the answer. Method-focused, because a system, like a paper, can reach agreeable results through a process that does not support them, and it is the process that generalizes to the next million decisions. Adversarial, because a review that is not trying to reject is not a review; it is a courtesy. And ongoing, because a decision system, unlike a published paper, keeps deciding — the method drifts, the inputs shift, and the referee's question has to be asked not once at admission but continuously against what the system is actually doing.

This is a different demand from the one science's own motto makes, and the difference is the point. To insist on inspecting the demonstration yourself — to take nobody's word for it — is to refuse an authority and check the thing directly. Refereeing is the institutional cousin of that instinct: not every relying party can perform the scrutiny, so we build a standing office of disinterested experts who perform it on everyone's behalf, before the claim is admitted rather than after harm forces the question. It is also distinct from the audit that arrives afterward. The referee examines the method before it is trusted; the auditor examines the record after it is relied upon. Both matter, and a serious accountability regime for machine decisions will want both. But the older discipline — probing how a claim was reached, by someone with no reason to spare it, as a condition of letting it into the record — is the one we have most completely dropped on the way from the laboratory to the model, and the one a world of fast, fluent, unrefereed machine decisions can least afford to skip. We already invented the answer. We are, once more, rediscovering it the expensive way.

— Dispatches · Summit Cognitive


Sources

  1. On the Philosophical Transactions (from 1665) as an early scientific journal and the Royal Society's assessment of submissions: Royal Society, "History of Philosophical Transactions"; "Philosophical Transactions of the Royal Society," Wikipedia.
  2. On the long, uneven development of refereeing and its systematization as a routine gate in the twentieth century: "Scholarly peer review," Wikipedia; Melinda Baldwin, "Scientific Autonomy, Public Accountability, and the Rise of 'Peer Review' in the Cold War United States," Isis 109 (2018).
  3. On peer review as a disciplined check rather than a guarantee — its real limits in catching fraud, error, and bias: R. Smith, "Peer review: a flawed process at the heart of science and journals," Journal of the Royal Society of Medicine/BMJ.

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.