The honest version of an agent capabilities chart, were anyone willing to publish it, would have two lines. One would show measured performance on benchmarks. The other would show something like the half-life of those benchmarks: the number of months between publication and saturation.
Both lines have been climbing. The first because models are genuinely improving. The second because the field has gotten extremely good at the specific game of beating evaluations, which is not quite the same game as the one we believed we were playing.
This is not a story about cheating, in the narrow sense. It is a story about Goodhart’s Law operating at a speed and intensity that classical economists would have found difficult to imagine. A benchmark, once published, ceases to measure capability and begins to measure attention. The labs that score highest are the labs that paid the closest attention.
Several research groups have begun arguing, persuasively, that the answer is not yet another benchmark but a different category of measurement entirely: held-out tasks generated dynamically, evaluations whose specifications are themselves agents, longitudinal studies of deployed systems rather than laboratory ones. None of these are easy. All of them are slower and more expensive than the current regime.
The slowness is the point. A measurement that cannot be saturated in six months is, almost by definition, a measurement of something more interesting than the saturation rate. Whether the field has the patience for such measurements, in a moment when patience is the scarcest resource of all, is unclear.
0 Comments