DISPATCHES · Summit Cognitive

← All dispatches

EvidenceSeptember 14, 20268 min read

How would we know?

An AI commitment needs evidence that can survive examination. So does an AI-assisted decision.

How would we know? The record has to outlast the room. An abstract document and an offset correction sheet, with the Summit Cognitive logo.

The decision is six months old. The model that produced the recommendation has been replaced. The reviewer who approved it has moved to another team. Now the applicant is back with one question: did anyone consider the correction I sent before you decided?

The case is hypothetical. The question is specific: can the institution establish what happened to that correction?

On paper, the institution looks prepared: a policy, a model evaluation, an approval, a log of the recommendation, and the correction itself, right there in the file.

But presence is not use. The correction’s presence does not establish that the system used it, that the reviewer saw it, or that it affected the decision. Those are three different claims.

Four separate claims, not a sequence. In the file: what establishes receipt and version? Used by the system: what establishes use in this run? Seen by the reviewer: what establishes what was presented? Affected the decision: what supports influence, and what remains uncertain?
Four claims. Each needs its own evidence. View the diagram at full size.
A file can contain the correction and still leave the applicant's question unanswered.

Put someone who was not in the room in front of that record. See how far they get.

That is where I would start the larger conversation now taking place about AI verification. In September, Dario Amodei proposed pacing frontier development. As a first step, Anthropic committed to embedded third-party evaluators with ongoing access and publication rights, subject to specified exceptions. Verification is part of his proposal. The work ahead is to establish how the arrangement operates, what its findings support, and where its limits lie. Amodei's proposal

A commitment, a way to check it, and evidence that the check works are three different achievements. We should give each its due without letting one stand in for the others.

A credible commitment needs evidence that can be examined, a competent examiner and the freedom to report an unwelcome finding. Institutions making consequential AI-assisted decisions face their own version of that requirement, even if frontier development slows. A laboratory can produce a valuable evaluation of a model. The institution still has to account for the decision that reached the applicant.

This is a familiar demand: account for a consequential decision. In the example here, the question crosses the source record, the model’s recommendation and the human review. Each stage can have a record while the connections between them remain unestablished. That is where I would look.

What the record owes the question

In this hypothetical, the institution may already have everything it needs. Historical source versions, the correction's arrival, the material shown to the reviewer and the policy in force may all be recoverable from ordinary systems. If those records answer the question at an acceptable cost, use them.

The test has to permit that result. An inquiry designed to discover a need for new software is a sales demonstration with the conclusion filled in.

Where the record falls short, name the gap. Today's database value alone cannot establish what was available six months ago. Today's policy, by itself, cannot establish which rule applied then. A note that a person approved the decision may say little about what that person could see or change.

The gap matters, but failure to establish that the correction was considered does not, by itself, establish that it was ignored. State what the available evidence supports and what remains unresolved. The next reviewer needs to know where the evidence stops.

I would ask four questions of the record.

What information was actually consulted? Its origin, version and condition matter, as does the reason for relying on it. An authentic document can contain a mistake. An integrity check can establish that captured material has not changed without establishing that its contents were true. A reviewer needs to know which claim has been checked.

An unchanged record can still be incomplete. What fell outside capture? Compare the account with the steps and records the workflow was expected to produce, and with other sources where available. An unexplained absence should remain visible. The institution’s chosen file cannot set the boundaries of the question.

Which rules and whose authority applied? That includes the policy, permitted exceptions, required review and authority to decide. If a required review did not occur, the record should preserve that discrepancy. Rewriting the account to resemble the intended procedure defeats the examination.

Can someone else reconstruct the material sequence? Give a competent reviewer appropriate access and enough independence to question the account. They need evidence of the inputs, outputs and human actions material to the question. Running today's model may be informative, but a new output cannot establish what the original reviewer saw. Nor does a difference between runs explain its own cause. Establishing that the correction reached the reviewer would answer one question. Establishing how it influenced the decision may require a different inquiry.

What can the affected person do with the answer? An internal audit may be satisfied while the applicant remains unable to pursue a correction. A useful review process gives that person relevant information and a route to correction. It should not require the applicant to reconstruct the institution’s process before anyone will examine the objection. Someone must be responsible for taking it up, including when the available record cannot settle it. Material that would expose someone else's information may need to remain with an authorized examiner. Those access boundaries should be explicit. The record can support an available review process; it cannot create legal standing on its own.

The answer has to remain useful after the people who made the decision have moved on. The record has to outlast the room. Recollection may help; it should not be the only way into the case.

The frontier can slow. The question remains.

Pace can be measured within a defined scope. Release dates, benchmark results and development inputs tell us different things. A trend on one family of tasks may not describe another. A credible commitment therefore needs agreed measurements, comparable conditions and access to the work behind the result.

The model and the decision are different objects of examination. A successful evaluator arrangement could tell us a great deal about a model and its development. It would not, by itself, tell us what happened to this applicant's correction. Both examinations matter. Each has to stay within what its evidence supports.

Slower capability growth would not withdraw decisions already made. Deployed systems may remain in use, and reviews, renewals or disputes may still require an account of their work. Replacing the model does not answer a question about the decision it helped make.

There is a temptation during rapid improvement to defer that work. Tomorrow's model will be better. Today's deployment will soon be obsolete. Why spend time examining something you expect to replace?

I think of that temptation as an acceleration amnesty: the prospect of improvement becoming a reason to postpone an account of present performance. This is a hypothesis about incentives, not a measured description of every AI program. A plateau could end the excuse. It could also leave less money for the work.

That is the commercial uncertainty we have to face. An evidence problem does not guarantee a budget. Deployment may shrink. Existing suppliers may satisfy the need. An institution may understand the problem and have more urgent uses for its money.

The question for a buyer is concrete: where does a missing or unusable record create work someone is prepared to pay to reduce?

Give the question a cost

Choose one consequential workflow and one review question. Have its owner and a reviewer responsible for challenging the answer agree, in advance, what would resolve the question and what would leave it unanswered. Include the objection an affected person would reasonably raise. Give a qualified colleague who did not make the decision the records they are permitted to examine. Watch the work.

How long does it take? How often do they have to call the original team? Can they recover the source versions, connect a rule to its effective date, identify an exception and distinguish the original record from an explanation written later? Which questions remain open? Whose work has the process reduced, and whose has it increased?

If the exercise exposes a gap, repair the smallest one that appears to matter. Then try a different permitted example, chosen by someone other than the workflow’s owner, with a reviewer who has not been coached through the first. Keep the agreed acceptance criteria. Include a case the existing process already handles well, so the improvement does not quietly make ordinary work harder. One example can reveal a problem. Its frequency and the value of a repair need their own evidence.

The result might justify a new capability. It might justify connecting records the institution already owns. It might show that the proposed change costs more than it saves. Each result gives the owner something to decide on.

The exercise must also make room for an uncomfortable answer. A technically intact file can preserve a bad decision. A complete chronology can expose an unfair policy. Evidence is useful because it gives a reviewer grounds to challenge the institution's preferred conclusion as well as support it.

That purpose should govern collection. Retain material for a defined review need, with appropriate access and retention. Keeping everything can create burdens and harms of its own. The measure is whether the record supports the examination it exists for.

I would keep three claims separate: what an institution owes the person affected, what the evidence can establish, and what a buyer will fund. If existing records consistently answer these questions at acceptable cost, the case for additional infrastructure shrinks. If the important questions cannot be answered through retained evidence, the method needs to change. If collection creates more burden than the review benefit justifies, the design is wrong. And if buyers will not commit time, appropriate access and money to closing a demonstrated gap, attention has not become a business.

The applicant does not need that business to exist before the institution examines its records. It can test its ability to answer now.

Six months from today, the model may have changed again. Someone else may have left. The policy may have been amended. What will remain for the person asking, and for the reviewer expected to answer?

How would we know?

— Dispatches · Summit Cognitive

Continue the conversation

Start with one review question.

If you own an AI-assisted workflow that is difficult to review, I would welcome a conversation about one question your records struggle to answer. Start with the workflow, the review burden and the person responsible. Leave the underlying case files in your systems. That is enough to discuss whether a bounded review would help.