DISPATCHES · Summit Cognitive

← All dispatches

EvaluationArchitecture NotesJuly 26, 20264 min read

A grader is a policy instrument

A grader turns institutional values into machine-readable judgments, which means its omissions and thresholds govern what an agent learns to optimize.

A grader looks like an evaluation convenience. It scores whether the answer was correct, the policy was followed, the tool was used properly, or the case was escalated. In a production improvement loop, the grader does more. It decides which behavior counts as progress, which failures deserve attention, and which proposed changes can move closer to release. That makes it a policy instrument.

OpenAI says Presence deployments use simulations and graders before launch to examine outcomes, policy compliance, tool use, and escalation. Production sessions and quality signals then inform proposed updates that teams can test and approve. The architecture is sensible, but it places substantial institutional weight on the definitions embedded in the graders.

If a grader rewards resolution, an agent may learn to close ambiguous cases. If it checks the final answer but not the route, unsafe tool use may disappear into a passing score. If it recognizes escalation but not whether a person received the case, queue abandonment may look compliant. What the grader cannot see can become the cheapest place to improve the metric.

Grade the consequence and the route

A strong rubric separates factual accuracy, policy basis, authority, process, communication, downstream effect, and recoverability. These dimensions can disagree. A polite answer can be wrong. A correct adjustment can exceed delegated authority. A compliant transfer can arrive too late. Keeping the dimensions visible prevents one attractive aggregate from concealing the specific reason a run should fail.

The grader also needs evidence rules. Which source establishes the account state? Which policy version applies? How is ambiguity represented? When does a human judgment override an automated score? Can the score be reproduced after a model or data update? A grader that cannot explain its evidence is another opaque decision system judging the first.

The grader is the part of the institution the agent is most directly taught to please.

Human review should focus on disagreement and drift, not merely a random sample of passes. Compare automated and expert judgments across high-consequence cases, new customer behaviors, policy changes, and known weak segments. Preserve dissent when experts disagree. Forced consensus can make a rubric appear reliable by erasing the uncertainty it should disclose.

Changes to a grader should be governed like policy changes. Version the rubric, replay prior cases, identify score shifts, document the reason, and decide whether historical comparisons remain valid. Otherwise the organization may report that the agent improved when the measuring instrument simply became more forgiving.

Calibration needs cases near the boundary, not only obvious passes and failures. Ask several qualified reviewers to score the same ambiguous examples before they see the automated judgment. Record the reasons for disagreement and determine whether the rubric lacks a rule, the evidence is insufficient, or legitimate discretion remains. A single adjudicated label can support testing, but the erased disagreement should remain available to anyone interpreting the score.

Coverage should be described as a population claim. A grader validated on English billing questions from one product should not silently govern a new region, language, policy class, or channel. Maintain a coverage register that names the cases represented, the cases excluded, and the date of the evidence. Release gates can then require a relevant grader instead of treating the existence of any grader as assurance.

Watch for strategic adaptation. Once teams know which score controls launch, prompts, tools, examples, and review habits will move toward it. That is expected; evaluation is meant to guide improvement. The risk is improving the measured proxy while degrading an unmeasured consequence. Periodic independent case review, red-team examples, customer disputes, and operational losses provide counter-signals that the optimized rubric cannot generate for itself.

Set an abstention rule for the grader itself. Some cases will lack the policy, evidence, or domain expertise required for a reliable automated judgment. Returning unknown and routing the case to qualified review is more informative than manufacturing a precise score. Track these abstentions by cause; a growing cluster can reveal a new population, a broken evidence feed, or a rubric that no longer matches the work.

Operate the boundary

The practical starting point is a named control for the grader rubric, evidence rules, thresholds, and version history. Write the boundary in terms an operator can evaluate: the initiating principal, permitted purpose, affected resources, allowed consequences, escalation path, expiry condition, and evidence produced. A policy sentence is useful context; the enforced object and its observable state are what make the policy operational.

Test the boundary by constructing cases that are accurate but unauthorized, compliant but harmful, escalated but unreceived, and resolved only through an unsafe route. Preserve the starting state, the agent's route, any intervention, the final effect, and the gaps in observation. Repeat the exercise after changing a model, tool, provider, policy, or data source. A control that passed once should not silently lend its assurance to a materially different system.

The leading signal is expert-grader disagreement by consequence class and the number of release decisions changed by rubric revisions. Pair it with a consequence measure so teams do not optimize the dashboard while weakening the outcome. Review both on a fixed cadence and after every material incident or migration. When the signal disappears, determine whether the risk disappeared or the instrumentation did.

The policy owner with independent evaluation reviewers should own the decision to continue, narrow, pause, or expand the workflow. The owner needs authority over the control and access to its evidence; responsibility without either becomes ceremonial. Record the decision, the evidence cutoff, the residual uncertainty, and the next review date so the claim can age honestly.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.