DISPATCHES · Summit Cognitive

← All dispatches

MethodThe CasebookJuly 27, 20265 min read

The exam the machine graded

A grade is a small verdict with a long reach — and when a model assigns it, the student meets a judgment about their mind that no one can quite explain and that will follow them into rooms they have not yet reached.

Consider a student — call her a senior in a large public program, one of several hundred sitting the same assessment. She writes an essay she is proud of, submits it through the portal, and two days later a number appears next to her name. It is lower than she expected, low enough to matter, and there is no comment attached, because no person wrote one. An automated grader assigned the mark. In the same window, a second message arrives: the remote-proctoring system has flagged her session for review. Her eyes left the screen too often. Her typing paused in a pattern the model found unusual. She is now asked, in polite institutional language, to account for behavior she did not know was being counted against her. She has met two judgments about her mind in one afternoon, and can see the basis for neither.

Education is where we make consequential judgments about developing people at scale, and it is the place where the method of judgment matters most and is examined least. A grade is not a description of a student; it is a claim about a student, issued by an institution, that travels. It rides along on transcripts, gates the next course, shapes what a scholarship committee or an admissions reader or an employer later believes before they have read a word the student wrote. We tolerate enormous opacity in how that claim is produced because we are used to trusting the grader. When the grader is a model, the trust does not transfer automatically, and the reasons it does not are worth stating plainly.

The proxy the model actually grades

An automated grader does not measure understanding. It cannot. Understanding is not a quantity available to a scoring function; it is an interior state we infer from outward signs. What the model measures is a proxy — some bundle of features it can detect that correlates, across the training data, with the marks human graders once gave. Essay length, lexical variety, the presence of certain connective phrases, structural regularities, similarity to high-scoring exemplars. The model rewards what it can see, not what the teacher meant. Most of the time the proxy and the target move together, which is exactly why the arrangement survives: it is right often enough to look like measurement.

But a proxy is not the thing, and the gap between them is not random noise — it is a channel that can be exploited and a channel that can betray. The student who learns the proxy beats the student who learns the subject. Write longer, hedge less, deploy the connectives the model favors, mirror the shape of the exemplar, and the number rises whether or not the thinking underneath has improved. This is not cheating; it is responding rationally to the thing that is actually graded. Over time an assessment that keys on a proxy teaches students to produce the proxy, and it quietly disadvantages the ones who wrote something true but unusual — the essay that is good in a way the model was never trained to recognize. The mark is a small verdict, but its reach is long, and one that rewards the wrong signal compounds every time it travels.

A machine that grades the proxy does not measure the education. It teaches the student to master the proxy, and then reports the mastery back as if it were the mind.

An accusation dressed as an observation

The proctoring flag is a different and sharper wrong, because it is not a measurement at all — it is an accusation wearing the clothes of an observation. A proctoring model does not witness cheating. It has never seen the student's screen the way an invigilator in a room can. What it detects is anomaly: gaze that departs from a baseline, timing that departs from a distribution, a face that leaves the frame, a second voice, a pattern the system has been tuned to treat as suspicious. Anomaly is a statistical fact. Guilt is a finding about a person. The system produces the first and the institution reads it as evidence of the second, and the distance between them is where the injury lives.

Anomaly is not guilt because anomaly has innocent explanations, and the innocent explanations are numerous, ordinary, and unequally distributed. A student looks away because she thinks with her eyes on the wall. She pauses because she is anxious, or ill, or nursing a child in the next room, or living with a condition that makes her stillness look like fidgeting to a classifier trained on someone else's. To be flagged is to be told that your body moved wrong while you were thinking, and then to be asked to prove a negative — to disprove a suspicion assembled from behavior you cannot see and were never told was load-bearing. There is a particular cruelty in doing this to the young. An integrity accusation is a claim about character, and to level it at a developing person on the strength of a classifier score, with the presumption running against them, is to teach a lesson far more durable than anything on the exam: that a machine's unease about you counts as an indictment, and that the burden of your own innocence is yours to carry.

What the account owes the student

The remedy is not a better model, though the models will get better and the flags more precise. Precision is not the issue; accountability is. A judgment this consequential owes the student a specific kind of account, and the account has three parts.

First, the actual basis of the mark or the flag, in a form a student and a teacher can examine. Not a number and not a euphemism — the features the grade rested on, the signal the flag fired on, rendered legibly enough that a person can look at it and say that is not what my essay was doing or that pause was me rereading the question. You cannot contest a conclusion handed to you without its grounds; contestability requires that the grounds be produced. Second, a human with the authority to overturn the machine — not a human who forwards the appeal back into the same system, but someone who can read the work, weigh the explanation, and set the determination aside. A right of appeal that cannot reach a decision-maker is not a right; it is a waiting room. Third, a record: a durable account of what the system determined, on what basis, and what a human did with it afterward. A grade travels and an integrity finding travels farther, and a wrong one, left unrecorded and uncorrected, hardens into a fact about a person that no later room will think to question. A Decision Receipt — the evidence the judgment rested on, the rules that were active, the outcome, preserved so it can be replayed and rebutted — is what lets a young person carry not just the verdict but the standing to challenge it.

The scenario above is illustrative — a composite drawn to show a pattern, not an account of any real person, company, or event.

None of this slows the machine down where it is right. Where the grade is fair and the flag is sound, the account costs the institution nothing: the record holds, the appeal confirms, the basis withstands a look. The account is expensive only where the judgment was weak — which is precisely where a student, standing at the start of a life the mark will follow her into, is owed the chance to fight it. A judgment about a developing mind, made at scale by a method no one can see, is not accountability until someone can be shown its reasons and someone with authority can be asked to answer for them.

— Dispatches · Summit Cognitive

Continue from here

Turn the argument into a practice.

Get new dispatches, assess how your organization handles consequential decisions, or explore Summit Cognitive.