For most of the last three years, the headline benchmark in artificial intelligence was a single, glamorous number: how well a model could answer one question, in one shot, in one sitting. It was a useful number, the way a sprinter’s hundred-metre time is useful. It also told you almost nothing about whether the runner could finish a marathon.
What has changed in the past nine months is that several research groups, working largely in parallel, finally began publishing models that could finish the marathon. They are not, by any aggressive definition, smarter than their predecessors on a per-token basis. What they have learned to do is keep going.
The trick, in retrospect, looks less like cognition and more like bureaucracy. Long-horizon agents now keep structured scratchpads, version their own plans, and re-read their own notes before acting. The most capable systems will revise a plan four or five times across a multi-day task, discarding entire branches when a sub-goal turns out to be unreachable. To watch one work is to watch a careful, slightly anxious project manager.
Researchers at three independent labs have published results in which agents complete software engineering tasks averaging fourteen hours of wall-clock time, with success rates above seventy percent on tasks that, eighteen months ago, no model could finish at all. The interesting figure is not the success rate. It is the duration.
Skeptics, and there are many, point out that wall-clock time is a strange axis on which to measure intelligence. A clock running for fourteen hours says nothing about the quality of the thought inside. They are right, and they are also missing the point. The point is that the thought, whatever its quality, did not collapse.
Whether this represents a path to something deeper, or merely a very expensive way of being persistent, is the question that will define the next year of research. The labs themselves seem unsure, which is itself a kind of progress.
0 Comments