The consensus of machines
When you ask three models and they agree, it feels like corroboration — but if they learned from the same data and share the same blind spot, their agreement is not three witnesses confirming a fact; it is one error, echoed three times, wearing the costume of consensus.
There is a pattern in agentic systems that feels like rigor and is often the opposite. You do not trust a single model with a consequential call, so you ask several — an ensemble, a jury of models, the same model sampled a dozen times and polled for its most frequent answer. When they converge, you relax. Three independent judgments pointing the same way; surely that is stronger than one. The design is meant to be humble, a hedge against any one model's error, and it wears the look of good epistemic hygiene. But the comfort it produces is doing more work than the method warrants, and the gap between the two is where trouble lives.
Agreement that is not corroboration
The intuition being exploited is old and mostly sound: if several observers, questioned apart, report the same thing, the thing is probably so. Two people describing the same car at an intersection tell you more than one, and ten tell you more than two. This is corroboration, and it is one of the load-bearing ideas in how humans decide what to believe. But the intuition carries a condition it usually leaves unspoken — the observers have to be separate. Their reports count as independent evidence only insofar as their errors are uncorrelated: only insofar as whatever might have led one of them astray did not lead all of them astray in the same direction.
Machine agreement quietly violates that condition. Models drawn from the same lineage are not separate observers. They are trained on overlapping corpora, shaped by similar architectures, tuned toward similar preferences, and drilled on much of the same public text. When you poll them, you are not canvassing three witnesses who saw the event from three positions. You are asking three copies of roughly the same reader what they took away from roughly the same library. Where the library is wrong, or silent, or slanted, all three inherit the flaw together. Their agreement then reports not that they each found the truth, but that they were each built to miss the same thing — and to miss it confidently, in unison.
This is correlated failure, and it is the default condition of the field, not an exotic edge case. The whole reason a second model is worth asking is that it might fail differently from the first. Two systems that fail in the same places give you one system with a louder voice. And because the polling machinery treats each response as a vote, the architecture is structurally blind to the distinction: it counts the ballots and cannot see that they were filled out by the same hand.
Unanimity as a stronger illusion
What makes this more than a technicality is that unanimity does not merely fail to add evidence — it actively adds persuasion. A lone answer invites doubt; you know it is one fallible source, so you keep some skepticism in reserve. A unanimous panel disarms exactly that reserve. Nobody dissented, so where would you even begin to object? The consensus arrives already having answered the obvious challenge, and its very completeness makes it feel less like an opinion and more like a reading off the world. The more total the agreement, the harder it is to find a foothold from which to push back.
So the failure mode is not that the models are sometimes wrong. Everything is sometimes wrong. The failure mode is that they can be wrong together, and the togetherness converts a shared error into something that behaves, socially and procedurally, like an established fact. A dissent would have been a gift — a crack to work at, a reason to look closer. Unanimity removes the crack. You are handed a verdict with no seam, and the absence of a seam is read as strength when it may only be the signature of a common cause.
Three models that learned from the same world will make the same mistake and call it agreement — and you will believe them precisely because they are unanimous.
The law worked this out a long time ago, in the humbler setting of human testimony. Witnesses corroborate only if they did not copy from one another. Two accounts that match are strong evidence when the witnesses were kept apart and weak-to-worthless when they rode to the courthouse together and compared notes on the way. Collusion, shared rumor, a single upstream source feeding both — any of these turns matching testimony from confirmation into echo. The doctrine is not squeamishness about agreement; it is a precise recognition that agreement is only informative in proportion to the independence of those agreeing. A machine ensemble trained on a common inheritance is the modern equivalent of witnesses who all read the same newspaper before they took the stand.
Recording independence, not just agreement
The corrective is not to stop polling models, which is often genuinely useful, but to stop treating the agreement itself as the finding. Machine consensus is evidence of consistency — that several related systems process an input the same way — and consistency is worth something. It is simply not the same thing as correctness, and the method's whole persuasive force comes from conflating them. What you actually want to know, before you let unanimity carry weight, is whether the sources that produced it were independent enough for their agreement to mean anything.
That is a question you can only answer if you record more than the vote. A verdict that logs three matching answers and a tally has thrown away the one fact that determines what the tally is worth. What belongs in the record is the lineage: which models, from which families, trained on what — enough of the sources' provenance to judge how much of their common answer traces to a common origin. Diversity is the property being assessed, and diversity is invisible in the output. Two answers look the same on the page whether they came from genuinely different systems or from the same system wearing two names. Only the provenance tells them apart, and only if someone thought to keep it.
Treat a consensus of clones for what it is: a single point of failure in disguise. The disguise is the whole problem — a solitary model that is wrong announces itself as one fallible source, while the same error spread across an ensemble presents as a chorus and is trusted accordingly. The point is not to distrust agreement but to price it correctly, and you cannot price it without knowing how independent the agreeing parties were. A decision record that captures the diversity of what produced a verdict lets a later reader do that arithmetic. A record that captures only the verdict asks them to mistake the size of the choir for the truth of the song. Corroboration was always about independence. When the witnesses are machines cut from the same cloth, that is the one thing their unanimity cannot vouch for — and the one thing worth writing down.
— Dispatches · Summit Cognitive
Continue from here
Turn the argument into a practice.
Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.