AI Workforce

METR 50% Task-Completion Horizon

Also tracked as The Horizon

A fitted 50 percent success threshold for METR benchmark tasks, not elapsed agent runtime

What is the current METR 50% Task-Completion Horizon reading?

METR 50% TIME HORIZON
17.4 hours
of human-expert task time at METR's 50% success threshold; not elapsed runtime
Apr 2025
2 hours
For a model released in April 2025. The latest reading is higher.

METR's latest 50% time horizon is 17.4 hours, for a model released in April 2026. That is the length of a task, timed by human experts, that the AI agent is predicted to complete half the time on METR's software-heavy tests. It is not how long the AI can run on its own, and METR says its estimates for the longest tasks are unreliable.

Measurement basis: METR's 50 percent time horizon is the estimated human-expert task duration at which the evaluated agent's fitted success probability is 50 percent; it is not elapsed agent runtime. The current task suite primarily covers self-contained, well-specified software engineering, machine learning, and cybersecurity tasks with clear success criteria, and METR says measurements above 16 hours are unreliable.

On METR's software-heavy tasks, a model released in April 2026 is predicted to succeed half the time at tasks that take a human expert 17.4 hours, the longest horizon among the models METR has measured on its current method.

METR, a research nonprofit, times how long human experts take on a set of tasks, then tests AI agents on the same tasks. For each model it fits a curve and reads off the task length at which the agent is predicted to succeed 50% of the time. For a model released in April 2026, that length was 17.4 hours of human-expert time. For a model released in April 2025, it was 2 hours.

The horizon measures how hard a task is, not how long the AI works on it. METR says that on the tasks agents do finish, they usually take a fraction of the time a person would. It also does not mean the model succeeds on half of all tasks of that length: on some tasks a model succeeds every time and on others it fails every time.

The estimates are rough. Error bars have been about a factor of two each way, and METR says values above 16 hours (960 minutes) are unreliable on its current tasks. METR also revises past values when it changes its method, so the line is best read as a long-run slope, not a step from one model to the next.

The tasks are mainly self-contained software engineering, machine learning and cybersecurity problems with clear pass-fail checks, done by people with little context. The results do not carry over to office work, computer use or jobs in general. A capability test is not a labor-market outcome: The business AI adoption rate tracks business use and the AI layoff count tracks announced layoffs, and none of these shows that one caused another.

Source: METR (Model Evaluation & Threat Research) · Source data ↗ · Latest: Apr 2026

Is this happening to you?

Does your work involve software, machine-learning or security tasks like these?

METR 50% Task-Completion Horizon over time: what has changed?

CSV Chart Card
METR 50% time horizon, model released April 2026: 17.4 hours
Hours of human-expert task time at 50% predicted success, by model release date; METR says values above 16 hours are unreliable
METR 50% Task-Completion Horizon
Historical data
Event-driven · METR (Model Evaluation & Threat Research)
Period Value YoY Change
Apr 2026 17.4 hours —
Feb 2026 12 hours —
Dec 2025 5.9 hours —
Nov 2025 4.9 hours —
Aug 2025 3.4 hours —
May 2025 1.7 hours —
Apr 2025 2 hours —
Feb 2025 1 hour —
Dec 2024 0.7 hours —
Oct 2024 0.3 hours —
Sep 2024 0.3 hours —
Jun 2024 0.2 hours —

Frequently Asked Questions

What is METR's time horizon?

It is the length of a task, measured by how long human experts take, at which METR predicts an AI agent will succeed half the time. The latest reading on this page is 17.4 hours, for a model released in April 2026. It measures task difficulty, not how long the AI works.

Does this mean AI can do a full day of someone's job?

No. METR's tasks are mainly self-contained software engineering, machine learning and cybersecurity problems with automatic scoring, done by experts with little context. METR says its human times likely overstate how long an expert takes on the job, and its separate study found much shorter horizons for visual computer-use tasks.

How reliable is the latest number?

The latest reading on this page is 17.4 hours. METR says values above 16 hours (960 minutes) are unreliable on its current tasks. METR's own error bars have been about a factor of two each way, and it revises past estimates when it changes its method.

Where does this data come from?

METR, a research nonprofit, publishes time horizons for the AI models it chooses to evaluate. It says its list is not a complete record of the most capable models, and each point on the chart is dated by the model's release, not by when METR measured it.

Ross Kilburn
Written by

Ross Kilburn, Founder

Former COO of Ark Law Group, a foreclosure defense firm serving five states · founder of Seattle Short Sales · author of Short Sale Your Home

Ross Kilburn is the former COO of Ark Law Group, a foreclosure defense firm serving five states. He founded Seattle Short Sales, wrote Short Sale Your Home, and worked as a mortgage loan originator and real estate agent. He founded American Default Research in 2026.

Read more
from Ross →

Quick poll

Is this affecting you or your household?

No name, contact details, or raw IP stored · IP-derived code and answer kept 30 days to prevent duplicate votes

Create a free account to save indicators to your watchlist and get weekly updates.

Create Free Account →

Discussion

Loading comments…

Sources and methodology

American Default Research tracks 105 live indicators of household financial distress, including this one. The methodology page explains where each comes from, how often it updates and how the index uses it.
View methodology →