Dispatches · Summit Cognitive

All dispatches

Decision assuranceFlagshipAugust 11, 20267 min read

The Output Is Not the Decision

Governance fails when it treats the model's answer, the institutional decision, and the resulting action as one event. Keeping them apart is the work.

In brief: An AI answer, an institutional decision, and the action that follows are different events. A useful decision record preserves the evidence for each one and states what it cannot prove.

Three events happen. Most governance programs collect evidence for one of them.

A model output is a statement.

A decision is an exercise of authority.

An action is a change in the world.

They can occur milliseconds apart. They are still three events, and they fail independently.

The distinction is easy to lose because agent systems can represent all three in a single flow. A model proposes an answer. An orchestration layer selects a tool. The tool submits a request. The receiving system changes state. The interface reports success. One transcript holds the sequence, so the sequence reads as one thing.

It is not one thing, and the failure modes do not line up.

The output can be accurate while the decision is unauthorized.

The decision can pass evaluation while execution fails.

The action can complete after the caller has timed out and reported uncertainty.

The record can verify cryptographically while the evidence inside it is false.

A sufficient decision receipt of any design should preserve those boundaries rather than erase them. It should identify the scope of the decision, the principal, the authority relied on and its limits, the evidence with provenance, the policy version in force, the verdict and its reasons, the integrity material, the retention and redaction rules, and the path to redress. It should state the execution boundary explicitly.

That last field prevents a common category error. A receipt can establish that an evaluation occurred and that the record of it has not been silently changed. It does not thereby establish that a downstream action executed, that the action produced its intended effect, or that the decision was substantively correct.

Those require separate evidence, gathered by different means. An institution that cannot name which evidence it holds does not know what it can defend.

Signed and hash-linked decision records exist outside Summit. Possession of the primitive is not the company thesis. The next question is sufficiency.

The Receipt Sufficiency Test

A relying party should be able to inspect ten fields in one record.

  1. Scope. Which exact decision is covered?
  2. Principal. Who or what acted?
  3. Authority. Under what grant, and with what limits?
  4. Evidence. What was relied on, and does it carry provenance and freshness?
  5. Policy. Which version was in force at evaluation time?
  6. Verdict. What disposition and reasons were recorded?
  7. Execution boundary. Was action attempted, completed, failed, or left unknown, and what separate evidence supports that state?
  8. Integrity. What do the signature and hash cover, and what do they not cover?
  9. Retention. How long does the record survive, and what redaction rules apply?
  10. Redress. How can the decision be appealed, reopened, or superseded?

Each field is inspected against a specific record. This is not a maturity model or a list of principles. It is a set of questions with an artifact in front of them.

A weak receipt can still be signed.

It can omit the evidence that decided the matter. It can carry a vague authority claim. It can bind to a policy identifier no reviewer can retrieve. It can report success without distinguishing submission from effect. It can become unreadable outside the platform that produced it.

The signature can protect a narrow integrity property. It is not a substitute for a complete argument.

A declared composite

The following scenario is assembled from common patterns. It is not a customer, incident, or deployment.

An internal review asks an institution to justify one consequential determination made several months earlier. The operator produces the record. It is signed. The signature verifies. Then the reviewer applies the ten fields.

The policy identifier resolves to the current policy, not the one in force at evaluation time. The evidence is listed as an internal dataset with no retrieval time. The principal is a service account with no accountable role behind it. The verdict reads success, which means the evaluation passed, though people in the review assumed it meant the external action completed. There is no redress field, and the affected party has already been told the matter is final.

Every stated integrity property held. The record still failed to answer the questions the review was convened to ask.

That gap is the subject of this essay.

The strongest objection

The serious counterargument is that signatures are objective while sufficiency is subjective. Requiring a relying party to judge the record might appear to replace a control with an opinion.

Three answers.

First, the subject of the test is a record, not a program. Each field is checked against a specific artifact and produces an inspectable answer. The relying party supplies the threshold. The record supplies the evidence.

Second, the relying party supplying the threshold is the point. An operator that defines the adequacy of its own evidence has performed a self-assessment and called it assurance. Evidentiary systems put the standard in the hands of the party expected to rely on the record.

Third, a signature is objective about a narrow property: whether the covered bytes changed after signing. Objectivity about one property is not coverage of every property. Precision is not sufficiency.

When Summit issues a receipt, Summit is part of the system under review. That makes the integrity field a question about us as well. We would rather a relying party ask it than assume the answer.

What the test does not prove

The test does not establish that the evidence was true. It does not establish that the policy was legitimate or lawfully applied. It does not establish that the decision was substantively correct. It does not establish that the action executed or produced its intended effect.

A record issued by the operator of the system under review is not, by that fact alone, independent attestation. Independence is a property of who can check, under what conditions, and with what materials. It has to be demonstrated rather than asserted, including by Summit.

Why the separation pays

Once output, decision, and action are held apart, governance becomes more precise.

Model evaluation assesses the output.

Decision assurance assesses the evidence, policy, and authority.

Execution telemetry establishes what the target system did.

Review and redress address the consequence for whoever bore it.

Different questions require different instruments and different evidence. A single green status is not a summary of them. It is a way of not asking them.

A system does not become trustworthy by compressing the events into one report. It becomes answerable by preserving the difference.

Take one decision your own systems made last week. Apply the ten fields. Count only what the record itself can support without asking the team that produced it.

The gaps are the finding.

Dispatches · Summit Cognitive