On a coding agent leaderboard, the score grabs your eye first.

But when your team actually runs an agent, you need at least three numbers side by side.

Success rate
Cost
Time

The Artificial Analysis Coding Agent Index publishes cost, token usage, and execution time alongside performance for exactly this reason.

Pass@1 Is Close to First-Try Success

Coding Agent Index v1.5 combines pass@1 from three evaluations: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA.

This metric matters because it does not measure the best result after unlimited retries. It measures the chance of solving the task in a single attempt.

But in a real product you can retry after a failure. So pass@1 alone cannot tell you the cost.

When 2 Points Higher Costs More

Imagine two agents.

Agent A
Score: 58
Cost/task: $4
Time/task: 12m

Agent B
Score: 56
Cost/task: $1.5
Time/task: 5m

Agent A ranks higher. Yet for a large volume of low-risk tasks, Agent B can be the better pick.

For a migration task where failure is very expensive, those 2 points can be worth it.

Context decides. The score alone does not.

CodeBridge Mini Lab: Draw the Efficiency Frontier

Run your three candidate agents on the same task set.

agent,success_rate,cost_per_task,time_min
A,0.82,3.8,11.2
B,0.78,1.4,5.3
C,0.65,0.3,2.1

Then plot success rate vs. cost and success rate vs. time.

Your goal is not to pick one winner. Your goal is to remove dominated options.

If one agent is more expensive, slower, and less successful, you have no reason to keep it.

Developer Time Is Also a Cost

A cheap API call can still be expensive. If the agent fails after 30 minutes, you must re-read the context yourself.

So real cost includes more items.

API cost
+ retry cost
+ human review time
+ failure recovery time

At a small company, 20 minutes of developer time can cost more than $1 of API usage.

You should track both.

A Fast Agent Can Change Your Workflow

Time is not just a UX metric.

If results arrive in 5 minutes, you can wait and review them right away. If they take 40 minutes, you need an async workflow with notifications.

In other words, latency shapes your product architecture.

Speed changes how you build, not just how you feel.

Conclusion: The Best Coding Agent Depends on Your Workload

Leaderboard scores are a good starting point. But for operations, ask three questions together.

How often does it succeed? How much does one attempt cost? How long until you get the result?

When you read those three axes together, your agent choice becomes far more realistic.

Further reading

References

Go deeper with a course

If you want to measure success, cost, and time on a real repository with Claude Code, guided practice helps.