AI’s Task Horizon Doubles Every Seven Months
The length of tasks AI can complete on its own, at 50 percent reliability, has been doubling roughly every seven months for six years. The measurement belongs to the independent research group METR: hundreds of real tasks were first timed on paid human experts, then handed to each generation of frontier models. The result is a curve that started at seconds, reached hours, and keeps compounding.
This number puts the most-skipped dimension of “which work to delegate” on the table: the answer is not fixed. Part of what is called “never” today belongs merely to “not yet” — and a business that confuses the two errs in both directions at once: carrying delegable work on human shoulders today, delegating undelegatable work tomorrow.
What the Number Says
BU BÖLÜMÜN ÖZETİ
- What is measured is task length, not intelligence
- Fifty percent is not a delegation threshold
- The curve was drawn on software tasks
Three details inside the seven-month doubling carry the decision value.
What is measured is task length, not intelligence
The study’s elegance is its unit: not how clever the model is, but how long a job — timed on skilled professionals — it can finish alone. That unit is the only one that translates into an executive’s language: “how many of my expert’s hours is this task” is the common currency of every delegation decision. The unit is honest, too — it leaves no room for hype or dismissal, because everyone knows the hours of their own work.
Fifty percent is not a delegation threshold
The headline curve is drawn at one-success-in-two. The same research measures higher reliability bars, where the horizon is naturally shorter. The business lesson is crisp: “can do” and “can be shipped” are different thresholds — in work where an error is cheap, even 50 percent earns its keep; where an error is dear, even 99 percent gets questioned. The threshold debate is therefore a work debate, not a model debate, and the work’s owner belongs at that table more than the technical team does.
The curve was drawn on software tasks
The measurement’s raw material is software and reasoning work — jobs whose output verifies quickly. In work that verifies slowly, carries scattered context, or defines “success” hazily, the horizon may run shorter. The number is not a universal calendar; it is a direction sign for verifiable work — and the first question any executive should carry home from it: does this task’s output check faster than it produces?
What the Headline Misses
BU BÖLÜMÜN ÖZETİ
- A moving ruler demands dated verdicts
- As tasks lengthen, oversight changes shape
- The compounding also prices the waiting strategy
The headline sells the compounding speed; the guidance hides in three corollaries.
A moving ruler demands dated verdicts
Against a capability that doubles every seven months, an undated delegation guide starts aging the day it is written. No “this task cannot be delegated” ruling should enter the books without its date and the tool generation it was judged against — the reason the delegability test ends with a date line is this curve.
As tasks lengthen, oversight changes shape
On a minute-long task, a human reads the output and approves; on an hour-long task, reading every step is no longer possible — oversight must migrate to sampling and checkpoints. The curve therefore rewrites not only “what can I delegate” but “how do I audit what I delegated”, every period.
The compounding also prices the waiting strategy
“Let it mature, then we start” forgets the curve’s other face: the rival is reading the same chart. Task definition, audit routines and delegation habit are cumulative assets; as the curve compounds, so does the distance between the prepared and the unprepared. Waiting looks riskless — yet against an exponential, waiting has a price too, and its invoice arrives silently.
What a Business Should Do
BU BÖLÜMÜN ÖZETİ
- Label tasks with hours
- Set the threshold by the error’s price
- Put the guide on a calendar
Three decisions fall out of the curve.
Label tasks with hours
The task inventory gains one column: how many minutes or hours of a competent person is this job? The label shows at a glance which tasks sit inside today’s horizon and which are approaching it — the delegation queue reads off this column. The labelling need not be precise; even a rough hour estimate puts tasks into a discussable order.
Set the threshold by the error’s price
Each task family gets its own reliability bar: cheap, reversible errors take a low bar plus fast checks; dear errors take a high bar plus a human-approval flow. Deriving the threshold from the work’s nature rather than the technology’s mood is this set’s master rule.
Put the guide on a calendar
The delegation guide is re-scored every six months: the “never” list, the “not yet” list and the “delegate” list refresh in one meeting. Carrying a fixed guide against this curve is looking for a bolted bridge across a flowing river.
The study itself is at METR’s publication.
The ruler lengthens every seven months; whoever rules without dating the verdict never notices they are measuring with an expired ruler.
