How fast is AI?
Menu

Task length · data updated Oct 2, 2026

When can AI do a task as long as yours?

METR measures how long a task, in skilled-human time, frontier AI agents can complete on their own. That length has doubled about every 4.2 months since 2023, and every 6.2 months since 2019. Pick a length to see where the trend would put it.

A task that takes a skilled person

1 work week, half the time

Due now

On the post-2023 trend this was due around Sep 2026, but no measured model has reached it yet. METR tests only some models, and its tasks run out near 16 hours, so this may simply not be measured yet.

The longest task measured half the time so far is 17h of human time, by Claude Mythos Preview (Apr 2026), a provisional estimate past the edge of METR's task suite. Reaching 1 work week takes 1.2 more doublings: about 5 months at the post-2023 pace of one doubling every 4.2 months.

  • Aug 2026 · Fast end of the post-2023 range, doubling every 3.4 months
  • Sep 2026 · Post-2023 pace, doubling every 4.2 months
  • Nov 2026 · 2019-onward average, doubling every 6.2 months

Source: Starts from the longest horizon METR has measured and extends it at METR's fitted doubling times: the post-2023 estimate with its confidence bounds, and the slower 2019-onward average. Work units: 8-hour days, 40-hour weeks, 160-hour months, 2,000-hour years. · Task-completion time horizons (Time Horizon 1.1)

The measured curve

State-of-the-art models (50% reliability) Provisional, above 16 hoursTrend fitted to these points: doubling every ≈6.2 months (METR's own estimates: ≈6.2 months since 2019, ≈4.2 since 2023)
ModelsGPT-23sdavinci-0029sGPT-3.5 Turbo Instruct36sGPT-44mGPT-4 Turbo (Nov 2023)4mGPT-4o7mClaude 3.5 Sonnet11mo1-preview20mClaude 3.5 Sonnet (Oct 2024)21mo139mClaude 3.7 Sonnet1ho32hGPT-53.4hGemini 3 Pro3.7hClaude Opus 4.54.9hGPT-5.25.9hClaude Opus 4.612hClaude Mythos Preview17h*

Source: p50 task-completion horizons of state-of-the-art models (METR Time Horizon 1.1). Points above 16 hours are provisional; the dashed line extends the fitted trend. · Task-completion time horizons (Time Horizon 1.1)

How to read this

Length means human time

A task's length is how long a skilled professional takes to do it without AI, not how long the model runs. Models often finish far faster.

Half the time is not reliable

The headline measure is the task length a model completes 50% of the time. At 80% success the horizon is several times shorter, which is why the calculator lets you switch.

Benchmark tasks are tidy

METR's tasks are mostly self-contained software, machine-learning, and cybersecurity problems with automatic checks. Real jobs add ambiguity, coordination, and accountability, the weak links.

A trend is not a law

The dates assume the doubling simply continues. It could slow as tasks get messier, or speed up as AI helps build AI. Estimates above 16 hours are provisional.

Doing a long task is not the same as doing a job. A job is a chain of tasks, and output waits on the links people still hold. See your job task by task or read why weak links set the pace.