The Exponential Growth of AI: Task Horizons Double Every 7 Months
METR found that AI task horizons doubled every seven months from 2019 to 2025. Learn what the metric measures and why it matters for business.
Artificial intelligence improves with each generation of models, but academic benchmarks reveal little about how much real work a system can handle. METR proposed a more practical measure: the time a skilled human would need to complete the tasks that a model can finish on its own.
A measure of real work
METR uses the term task horizon for the duration of tasks that a model completes at a given success rate. When researchers ordered results from 2019 to 2025, they found that this horizon had doubled about every seven months.
The trend does not mean that every model improves at the same pace or that any task can be automated. It describes progress within the study’s benchmark, which focuses mainly on software tasks performed on a computer.
Why a model can pass an exam and fail at a project
Models can solve knowledge problems, write code, and outperform specialists on some tests. Their reliability drops when a task requires them to maintain context, use several tools, and correct mistakes over several hours.
The study shows the decline:
- On tasks that take a person less than 4 minutes, the best models approached a 100% success rate.
- On tasks that take more than 4 hours, the success rate fell below 10%.

Human completion time therefore works as an approximate measure of difficulty. Longer tasks give the model more opportunities to lose information, choose the wrong tool, or follow a path that does not produce the required result.
What a one-hour task horizon means
In the evaluation published in 2025, the strongest models tested, including Claude 3.7 Sonnet, reached a task horizon of about one hour at a 50% success threshold. In other words, they completed half of the tasks that would take a skilled person around one hour.
The figure changes when you require greater reliability. A 50% success rate is too low for unsupervised automation; at stricter thresholds, the horizon falls to a few minutes. This gap explains why a model can impress in a demonstration yet still require human review in a production workflow.
What the trend allows us to project
If the pace measured from 2019 to 2025 continued through the end of the decade, leading systems could tackle projects that currently require several weeks of human work. The authors estimate that even a tenfold measurement error would shift the forecast by about two years because of the trend’s slope.
This projection remains conditional. The benchmark covers a specific class of tasks and does not capture everything a business project requires, such as negotiating priorities, interpreting ambiguous information, coordinating people, or accepting responsibility for a decision.
How to use this measure in your business
- Start with bounded tasks that have clear inputs, outputs, and quality criteria.
- Measure the current work: time spent, errors, and cost before introducing AI.
- Run a pilot with human review until you know where the system fails.
- Reassess capabilities every few months because technical limits change faster than many internal processes.
The task horizon provides a concrete way to track AI progress. For a business, the useful decision is to identify which tasks already fit within that horizon and which still need human context, judgement, and accountability.
Based on the research “Measuring AI Ability to Complete Long Tasks”
Do you want to identify which tasks you can automate with current technology? Book a free 30-minute session.