The Quiet Revolution of Long-Horizon Agents
Models that once forgot a conversation by lunch are now finishing week-long projects. The story is less about raw intelligence and more about memory, planning, and patience.
Models that once forgot a conversation by lunch are now finishing week-long projects. The story is less about raw intelligence and more about memory, planning, and patience.
An agent is only as useful as the verbs it can perform. A field once dismissed as plumbing has quietly become the discipline that decides which agents work and which do not.
Every new agent capability is announced alongside a new evaluation that, suspiciously, the announcing lab tops. The structural issue is not dishonesty. It is that the thing we want to measure does not sit still.
Honesty in language models was, until recently, treated as a matter of training data. The new generation of work suggests it is closer to a matter of architecture.
While public attention follows the frontier, a quieter wave of deployment is reshaping how routine work gets done. The story is less dramatic than the demos and considerably more consequential.
A year of daily use has changed how I work. It has not changed it in the directions I expected, and the parts that have changed are not always the parts I would have chosen.