DISPATCHES · Summit Cognitive

← All dispatches

GovernanceThe Field ManualJuly 27, 20265 min read

Put an expiry date on the model

A model does not fail on a date you can see; it decays quietly as the world it learned drifts out from under it — so decide when it must be re-examined before you deploy it, or it will keep deciding, with growing confidence, long after it stopped being right.

Go through your production models and find the one that has run longest without anyone asking whether it is still valid. That model is making real decisions today against a picture of the world it built at a moment now well behind it. Nobody deployed it as a permanent fixture on purpose; it simply never acquired a date on which someone was obligated to look. This is the default failure of deployed intelligence, and it is entirely preventable: give every model that makes consequential decisions an explicit review-or-retire date, set and recorded the day it ships, with a named owner accountable for the review. Not a hope that somebody notices when it goes wrong. A governed expiry.

The reason this matters is that a model does not age the way a bridge ages, with visible cracks you can inspect. It ages by mismatch. It learned a relationship between inputs and outcomes at a particular time, in a particular population, under particular conditions — and then it is deployed into a world that will not hold still. Populations change. Behavior adapts, sometimes specifically in response to the model itself. Upstream data shifts as the systems feeding it are rebuilt. The relationship the model captured weakens, slowly, until the map no longer matches the territory. And through all of it the model keeps producing outputs in the same confident register, because confidence is a property of its arithmetic, not of its continued correctness. Nothing about a stale model looks stale from the outside.

Models are shipped as if they were permanent

Watch how a model actually enters production and you will see the assumption of permanence baked in at every step. It is validated against data that was current when it was built, it clears its checks, it is deployed, and attention moves to the next thing. The launch is treated as the end of the work rather than the start of a clock. There is a moment where someone asks “is this good enough to ship” — and almost never a scheduled second moment where someone asks “is this still good enough to keep running.” The first question is asked up front, when the model and the world are aligned. The second, the one that actually protects you, tends to get asked only after a visible failure has already forced it.

That timing is the whole problem. The most expensive moment to ask whether a model is still valid is after it has produced a wrong decision consequential enough to notice — because by then the invalidity has been operating for however long it took to surface, quietly mispricing risk or misrouting people the entire time. Drift does not announce itself; it accumulates below the threshold of attention and crosses into visibility only once it has already done its damage. A model with no scheduled re-examination will keep enforcing an outdated judgment, with undiminished confidence, until reality embarrasses it in public. The confidence outlives the validity, and the gap between the two is where the harm lives.

An expiry date is a decision made in advance

Consider how the physical world handles this. Perishable goods carry an expiry date — and the date is not a prediction that the good fails on that day; the milk is very likely fine the morning after. The date exists because someone decided, in advance, when the good must be checked rather than assumed. That is exactly the discipline a model needs, because a model is a perishable good that behaves like a permanent one: it degrades on a schedule the world sets, but gives you none of the sensory cues — no smell, no visible spoilage — that would prompt anyone to check. So you supply the cue yourself, deliberately, by writing down when the checking is due.

Concretely, three things get recorded at deployment. A review-or-retire date, tied to the real drift rate of the domain and not to a round number on a calendar — a fraud model in an adapting adversarial environment perishes far faster than a model of something slow and structural, and the date should reflect that difference honestly rather than defaulting to “annually” because annually is tidy. A named owner, a specific person accountable for the review actually happening, because a responsibility assigned to a team is a responsibility assigned to no one. And the monitoring that would trigger an early review — the drift signals whose job is not to be watched but to move the date forward when the world moves faster than you guessed. The expiry is the committed backstop; the monitoring is the early-warning line that can pull it in.

A model with no expiry date does not last forever; it just fails without anyone having agreed, in advance, to look — and confidence is the last thing to decay.

Govern the expiry, do not just declare it

A date recorded and then ignored is worse than no date, because it launders inattention as diligence. So the governing move is to change the status of a particular state: a model in production past its review date without a recorded re-validation is a defect, not a normal condition. Not a nagging reminder someone can dismiss, not a backlog item that ages gracefully — a defect, in the same category as a failing test or an expired certificate, understood to be wrong until it is resolved. The instant the state carries that weight, the review stops being optional — a defect is something an organization must clear, not something it is free to notice at leisure.

This is also where drift monitoring earns its keep or fails to. Many teams answer the expiry argument with “we already monitor for drift” — and monitoring is genuinely useful, but monitoring without a committed decision point is watching a gauge that no one is required to act on. A dashboard showing a metric slide is not governance; it is scenery, until it is wired to a moment where someone must decide to revalidate, retrain, or retire. The expiry date supplies that moment. It converts the gauge from something observed into something that triggers an action, by guaranteeing there is a date on which the question gets asked whether or not anyone was watching the chart that day.

And to the objection that revalidation and retraining cost real money and attention: they do, which is precisely why the decision to spend them should be scheduled rather than left to accident. The cost of a governed review is knowable and can be planned; the cost of an ungoverned failure is neither. To the objection that some domains are stable enough not to need this — fine, then the expiry is long. A slow-changing model might carry a review date years out. But it still carries one, deliberately set and recorded, because the claim “this domain is stable” is itself a judgment that deserves a date on which it is re-examined. The point was never a short expiry. The point is a chosen one.

This is the same discipline a threshold demands, and for the same reason. An operating point gets an owner and a review date because a cutoff left unexamined becomes a verdict nobody chose still being handed down; a model needs exactly that, because a stale model is a policy no one chose still being enforced. It encodes a view of the world — who is likely to repay, what a pattern means — and that view, once the world has moved past it, is not neutral machinery quietly running. It is a standing decision that has lost its justification and kept its authority. Put the date on it while you still remember why you trusted it, name the person who has to look, and make the overdue state a defect. Do that, and the model ages the way anything responsible ages: on a schedule you set, checked by someone accountable, and retired when it should be.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.