
The conversation
On 19 March 2025, the nonprofit AI evaluation group METR published a blog post alongside a paper, Measuring AI Ability to Complete Long Software Tasks, submitted the day before. The paper proposes a single metric for comparing an AI agent's capability to a person's: the length of task, measured in the time a human with relevant expertise typically needs, that a model can complete with a 50% success rate, which the authors call the 50%-task-completion time horizon.
What the documents show
METR timed human experts completing tasks drawn from its own RE-Bench and HCAST suites plus 66 additional shorter tasks the researchers wrote for this study, then compared those human completion times against model success rates on the same tasks. The paper reports that a frontier model available at publication, Claude 3.7 Sonnet, had a measured 50% time horizon of about 50 minutes of human-expert time. Looking across model releases from 2019 onward, the paper reports the time horizon 'has been doubling approximately every 7 months', though it notes the trend 'may have accelerated' in 2024. The authors attribute the increase mainly to models becoming more reliable and better able to adapt after a mistake, alongside improved reasoning and tool use, rather than to any single capability jump.
The system boundary
METR frames the metric as a translation device between two different scales, not as a claim about a model's autonomy inside a live deployment. The measured time horizon describes performance on METR's own bounded task set, timed under METR's own conditions for human comparison; it says nothing by itself about a model operating inside a product with its own guardrails, human-in-the-loop review, or escalation path. The paper explicitly discusses the 'degree of external validity' of its results as an open limitation, meaning the reported minutes do not automatically describe performance on a task outside the study's own suite.
Where it fails
The paper cautions that its own extrapolation, projecting that within five years models could automate software tasks currently taking a person a month, depends on the trend continuing in the same shape it has followed since 2019, a projection the authors label uncertain rather than settled. It also flags that measurement choices, such as which tasks and which human timers were used, affect the absolute time-horizon numbers even where the overall exponential shape looks robust across checks. An operator citing this benchmark should be clear about which models and task types the underlying paper actually evaluated.
- Does the cited time horizon come from METR's own task suite, or from a different benchmark being described in the same terms?
- What is the trend's stated margin of error, and does the operator's use case resemble the tasks METR timed?
- Has the specific model in question been evaluated by METR at all, or is its time horizon an assumption drawn from a related model?
METR's own paper treats the time-horizon trend as the strongest available signal of its kind while stating plainly that it is not a certainty about any single future model.
Sources & reading trail
METR's own announcement of the time-horizon metric, its task set, and the doubling-time trend it reports.
Source published: 19 March 2025 · Retrieved: 16 September 2026
The paper's own methodology, the Claude 3.7 Sonnet time-horizon figure, and its stated limitations on external validity.
Source published: 18 March 2025 · Retrieved: 16 September 2026
Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.