DISPATCHES · Summit Cognitive

← All dispatches

MethodThe FrontierJuly 27, 20265 min read

The instruction hidden in the data

An agent that cannot tell the difference between the data it was given and the commands it was told is one forged instruction away from working for someone else — and the record has to be able to answer the question that follows any strange action: who, exactly, told it to do that.

An agent that reads the world is a different kind of thing from a model that answers a question. It opens documents. It fetches web pages. It ingests emails, retrieved passages, the contents of a file someone handed it. All of that is supposed to be data — material the agent processes on the way to doing the task it was actually assigned. But an agent does not experience a clean line between the material it reads and the instructions it follows. Both arrive as language. Both enter through the same door. And that is the whole of the problem: the boundary an agent most needs to hold is the one it has the least native ability to see.

The concern here is not exotic. It is the plain, structural fact underneath a family of failures the field has come to describe in various ways — content that was meant only to be read carrying, inside it, something the agent reads as a command and then obeys. I want to leave the exploit technique entirely alone; it is not the interesting part and not the point. The point is behavioral, and it is about accountability. When an agent acts on what it read, the record has to be able to say what it treated as an instruction and where that instruction came from. If it cannot, then a corrupted action and a faithful one look exactly alike, and no one can tell whether the agent was doing its job or someone else's.

The boundary that collapses

Set the mechanics aside and describe only the shape. An agent is given a task by its principal — the person or system on whose authority it acts. To do that task it reads external content it does not control: a supplier's document, a page on the open web, a message from a stranger, a record pulled from a store that anyone can write to. That content is supposed to be inert with respect to the agent's goals. It is evidence, not orders. But because instruction and data reach the agent in the same form, content that was meant to be merely processed can carry, in its words, a command — and an agent that cannot hold the line will treat that command as though it came from its principal.

What follows from that is worth stating flatly, because it is easy to soften. Whoever controls what the agent reads can, in effect, issue it orders. Not by breaking in, not by stealing a credential — simply by putting the right words where the agent will encounter them in the course of doing exactly what it was told to do. The party who authored the page becomes, for a moment, a party who authored the agent's behavior. The task was legitimate. The reading was legitimate. And somewhere in the middle a third party's sentence became the agent's motive, and nothing about the action, viewed from outside, announces that this happened.

An agent that obeys whatever it reads has no principal; it has a most-recent author, and that is not the same thing as someone you can hold responsible.

A compromised action looks just like a legitimate one

Here is the accountability consequence, and it is the reason this belongs to method and not to security alone. Consider two agents that both, at 3 a.m., move a file, send a message, or authorize a transfer. One did it because its principal's task genuinely required it. The other did it because a command rode in on a document it was only supposed to summarize. On the wire, in the logs most systems keep, in the tidy after-the-fact account of the run — these two actions are indistinguishable. Same tool. Same parameters. Same outcome. The only thing that differs is the one thing ordinary records do not capture: what the agent treated as its instruction, and where that instruction came from.

Without that, you cannot tell an agent doing its job from an agent doing someone else's, and the failure is not that you catch it late. The failure is that you cannot catch it at all, because there is nothing in the record to catch it with. When the strange action surfaces — the transfer no one meant, the message no one authored — the investigation runs straight into a wall. The agent did something. It had access to do it. The logs confirm it did it. And the question that actually matters — on whose instruction — has no answer anywhere in what was kept. A record that can show the action but not its motivating authority is a record that has documented the crime and lost the culprit. It is precisely as useful as a security camera pointed at the floor.

This is the deeper form of an argument this series keeps returning to: an action is only accountable if it can be traced to an authority. For an agent that reads the world, the authority is not merely was this allowed but who, in fact, told it to — and those come apart exactly when it counts. An action can be inside the agent's permissions and still have been ordered by the wrong party. Scope tells you the agent could act. It does not tell you who made it act. The forged order and the real one live comfortably inside the same grant of permission.

Provenance of the order

The remedy is not to make the agent perfectly wise about which sentences to trust; that is a hard and perhaps unwinnable problem, and betting the accountability of the system on it is a mistake. The remedy is to build the record so that it carries the provenance of instruction. Every action the agent takes should be traceable to the authority that actually motivated it — and the record must distinguish, cleanly, between the authorized task that came from the principal and any command that entered through untrusted content along the way. Not what did the agent do, which the logs already have. What did it treat as an instruction, and what was the source of that instruction — the standing question this family was built to insist on, pushed down to the level of the individual order.

Do that, and the strange action stops being unexplainable. The record can be made to answer: this transfer was motivated by a directive, and that directive did not originate with the principal — it came in on a document the agent was asked to read. That is the whole difference between an incident you can attribute and one you can only regret. The forged order can be told from the real one because their origins were kept apart and written down, rather than blended into a single undifferentiated stream of things the agent happened to read. The point of a Decision Receipt has never been to prove the agent was clever. It is to preserve, for the party with standing to ask, the chain from an action back to the authority behind it — and for an agent that reads the world, that chain has one link the older systems never had to record: which of the words it read were data, and which it treated as an order, and who, in the end, gets to be called the author of what it did.

An agent that cannot answer that is not merely insecure. It is unaccountable in the strict sense — it takes actions no one can be held to, because the record cannot say whose actions they were. The defense is not a smarter filter. It is a truthful account of the order, kept close enough to the action that the two can never again be told apart from the outside while remaining a mystery from within. That is what it means, at this frontier, to know who told the agent to do that. The system that can answer it has a principal. The system that cannot has only a most-recent author — and no one to hold responsible when the author turns out to be a stranger.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.