When people explain how far AI agents have come, there is a more intuitive question than benchmark score 65.

Work that takes a human 30 minutes, 2 hours, or 8 hours — how far can an AI agent finish alone?

METR's Task-Completion Time Horizon is one way to approach that question.

A time horizon is not continuous runtime

Clear the biggest misunderstanding first.

A 50% time horizon of 2 hours does not mean the AI runs nonstop for 2 hours.

METR sorts each task by how long it takes a human expert, then connects that to the agent's success rate.

A 50% time horizon roughly means:

the point where the agent is estimated to succeed 50% of the time on tasks that take a human expert about that long.

The 80% horizon demands higher reliability the same way.

Why express it in time?

Reducing task difficulty to one number is hard.

Fixing one line of a bug
vs
Implementing a library feature
vs
Investigating a system problem across many files

All three are coding tasks, but their complexity differs.

Using human expert time as a proxy lets you see one direction clearly: "can it handle longer and longer work?"

50% and 80% are completely different operating bars

A 50% success rate is useful for tracking research trends, but it can be far too low for real automation.

50% success
→ fails one time out of two

80% success
→ fails one time out of five

For payments, deployments, or data deletion, even 80% may not be enough when failure is expensive.

So when you read a time horizon, look at the required reliability together with the time, not the time alone.

CodeBridge Mini Lab: write down the human time of your own work

List 10 tasks you hand to AI, and record how long each takes you directly.

task,human_minutes,agent_success
rename_api_field,15,1
fix_small_test,25,1
investigate_memory_leak,180,0
write_release_note,40,1

Then watch how the failure rate changes as task length grows.

The goal is not to reproduce METR's numbers. It is to find where the autonomy boundary sits in your own work.

Long tasks fail for more reasons than intelligence

Long tasks have many steps, so small errors compound.

Wrong assumption
→ wrong file edited
→ test result misread
→ next step goes wrong too

That is why long work depends on more than the model:

  • progress tracking
  • context management
  • test and verification
  • checkpoints
  • retries
  • human approval

In other words, time horizon naturally connects to the harness problem.

Do not over-read the time horizon

METR measures time horizon mostly with software-related tasks, including RE-Bench and HCAST suites. It does not represent every job or every real-world task.

New models are also not always measured immediately. METR states plainly that some public releases may be measured late or skipped.

So do not generalize into "AI now replaces N hours of human work."

The more accurate statement is:

On a specific evaluation task distribution, at a specific success threshold, we observe a trend of handling longer tasks.

Conclusion: real agent progress shows in longer work, not in one good answer

Model differences on short questions can keep shrinking. But real work requires chaining many steps.

So when you evaluate an agent, ask this alongside "what is its score?"

Up to what length of work can I hand over reliably?

Once you see that boundary, design anything longer with human checkpoints in the middle.

Further reading

References

Go deeper with a course

If you want to turn horizon thinking into checkpoints, verification loops, and recovery design, a guided course walks through the full harness pattern.