Perspective · 01
What an Evaluation Measures
Why a system that performs well in demonstration can fail in the ways that matter.
A system that answers well in a demonstration has answered well on the examples chosen by the person demonstrating it. That is not a criticism of demonstrations; it is a description of what they are for. The difficulty begins when the demonstration is treated as evidence about the system rather than evidence about the examples, and the distinction is easy to lose because nothing in the room looks selective. The questions are reasonable. The answers are good. The conclusion drawn is that the system works.
What has actually been observed is capability under favourable conditions. What is needed before deployment is an account of behaviour under unfavourable ones. These are not the same measurement taken at different levels of rigour. They are different measurements, and only one of them has been taken.
The set nobody chose
An evaluation is only as informative as the set it runs against, and the set is the part of the exercise that attracts the least scrutiny. A set assembled by the team that built the system inherits that team’s assumptions about what the system is for. It contains the cases that were in mind during design — which are precisely the cases the design already handles.
The useful set is the one nobody chose: inputs drawn from the operating record rather than from imagination. The malformed, the truncated, the ambiguous, the ones where two answers are defensible and the difference is consequential. These are tedious to collect and awkward to score, which is why they tend to be collected after the first failure rather than before it.
A second-order effect deserves naming. Once a set exists it becomes a target. Work concentrates on the cases being measured, and performance on them improves faster than performance in general. A set that is never refreshed stops being a measurement and becomes a specification — one describing what the system was tuned toward rather than what it will meet.
Measuring the wrong thing precisely
A number does not become meaningful by being precise. A score aggregates over a distribution of cases that will not match the distribution in use, and the aggregate conceals the structure that matters: which kinds of input fail, whether the failures cluster, and whether they are the expensive ones.
A system can be right nine times in ten and remain unusable if the tenth case is where the consequence lives. Cost is rarely uniform across errors. Being wrong in a way that is visible and quickly corrected is a different event from being wrong in a way that is plausible, confident, and downstream of the point where anyone would think to check.
An aggregate is therefore a starting point rather than a conclusion. The question it cannot answer is what happens when the system is wrong — whether the error surfaces at all, who encounters it, and what it costs before it is caught.
The boundary is part of the system
Discussion of these systems concentrates on capability, which is the part that improves on its own. What usually decides whether a system can be deployed is the boundary: what it is permitted to do, what it must decline, what requires a human decision, and what happens when it is uncertain.
A boundary is a design artefact, not a disclaimer. It is expressed in the interfaces the system is given, the actions it can take without confirmation, and the point at which control returns to a person. A system with wide capability and no boundary is not more useful than one with a narrow boundary drawn deliberately. It is less deployable, because its failure modes are unbounded and therefore cannot be priced.
The practical test is whether the boundary can be stated as a claim someone would put their name to: this system does these things, does not do those, and escalates under these conditions. A boundary that cannot be written down has not been designed. The system has whatever boundary its integration happened to give it.
What a record is for
The component that gets deferred is the record of what the system did. Not logs of requests, but a reconstruction of a particular decision: the input, the version, the context retrieved, the output, and what followed.
The reason is not compliance. It is that questions about these systems arrive after the fact and in a specific form — why did it do this, on this occasion. A system that cannot answer that can be defended only in general terms, and general defences are unpersuasive against a specific complaint.
The record is also the only mechanism by which the evaluation set improves. Real failures are the material a useful set is made of. A system that does not retain them has no way of learning what it is bad at, and will keep being measured against the cases it was already good at.
Four components, one system
A model is not a system. A system is a model, a set that measures it against conditions it did not choose, a boundary on what it may do, and a record of what it did. These are not stages, and the last three are not refinements of the first.
The order in which they are usually built — model, then boundary, then record, then evaluation — is close to the reverse of the order in which they matter. The evaluation determines whether the model is worth deploying. The boundary determines whether it can be. The record determines whether the answer survives its first serious question.
What cannot be evaluated cannot be deployed responsibly, because there is no basis for the decision beyond an impression formed in a demonstration. What cannot be bounded should not be, because failure modes that are not knowable in advance cannot be accepted in advance. Those two sentences are the position. Everything above is the reason for holding it.