"We use coding agents too." That sentence is no longer a flex. The follow-up arrives immediately.
"So how many tasks reached production per $1 of agent spend?"
That is the backdrop for GitLab's Duo Agent Platform Impact Analytics in early access. It connects credit consumption by team, task, and model to real production outcomes instead of raw usage counts. In the same breath, GitLab reports agentic development active users up 200% year over year in the last three months. When more teams use agents, tooling must answer "are we using them well."
What token KPIs miss
Many teams still track one of these three.
Usage metrics: call counts, active users, adoption rate
Token metrics: input and output tokens, credits burned
Score metrics: benchmark pass@1, leaderboard ranks
None of them says whether the work finished. Tokens burned means work was assigned, not completed. A high score means passing exams, not surviving in your repo without reverts.
Two missing links:
- Reverts: a merged-then-reverted change counts in usage but subtracts from outcomes.
- Human hands: if an agent spends $1 and a human spends 2 hours fixing it, the true cost is not $1.
That is why KPIs must move from tokens to outcomes.
What Impact Analytics tries to do
Based on GitLab Docs and release announcements, the Impact Analytics and Impact dashboard direction looks like this.
Left: cost (credit consumption by team, task, model; Credits usage visibility)
+
Right: outcomes (status of Duo-created MRs, merges, deploys, DORA metrics, cycle time)
=
Question: "which spend connected to production?"
The 19.4 "Credits usage visibility GA" covers the left side, and the Impact dashboards attach the right side. Surfaces like the Data Analyst Agent that answer "how long do our MRs sit in review" in natural language belong to the same shift: from "how much AI did we use" to "did usage ship faster."
The Forrester TEI numbers read in the same language: 400% ROI, $7.5M three-year NPV, payback under six months, 80% faster onboarding, migration compressed from eight months to two, 40% QA and security-remediation savings, 20% individual productivity gains. It is vendor-commissioned research, so discount it, but notice the units are time, cost, and delivery rather than tokens.
A four-metric outcome set
No enterprise dashboard is needed. Solo builders can start with these four numbers.
| KPI | Definition | How to measure |
|---|---|---|
| AI cost | Total AI spend in the period | API and credit bills plus apportioned subscriptions |
| Completed tasks | Tasks surviving in production | Merged and still standing after 7 days |
| Reverted changes | Revert share | Revert MRs ÷ agent-attributed MRs |
| Human correction time | Time humans spent fixing | Sum of review and fix minutes (rough is fine) |
The headline derivative:
cost per success = AI cost ÷ completed tasks
true cost per success = (AI cost + monetized human time) ÷ completed tasks
revert rate = reverted ÷ agent-attributed MRs
The skeleton matches the cost-per-successful-task method. The denominator just gets stricter: from "judged a success" to "survived a week in production." Passing tests and surviving a week after deploy are different difficulties.
Split by model and task type for extra value. Run the same 10 tasks on models A and B and compare the four numbers. You will see "which model finishes our work cheaply" instead of "which model is smarter." For reading leaderboards, see the pass@1, cost, and time post.
CodeBridge Mini Lab: a two-week measurement sprint
Window: 2 weeks, scope: everything you assign to agents
One line per task:
- date / task / model / AI cost / success (Y/N, bar: merged + 7-day survival)
- reverted? / human fix minutes / failure type
After two weeks, aggregate:
- Total AI cost: __
- Completed: __ → cost per success __
- Revert rate: __%
- Total human fix time: __h → true cost per success with hourly conversion __
Decision:
- Revert rate above 20% means a workflow problem, not a model problem (fix review, tests, permissions first)
- If human fix time dwarfs AI cost, the "cheap model" is the expensive pick
- Lock the cheapest-per-success combo in as the default for the next two weeks
The log itself becomes portfolio material. Not "I tried AI" but "I tracked AI cost, completed tasks, reverted changes, and human correction time, then changed the workflow" — an engineering story that travels into hiring and partnerships.
Conclusion: measurement changes the workflow
The last line is short.
What gets measured in tokens grows in tokens. What gets measured in outcomes grows in outcomes.
Track the four numbers even on a personal project: AI cost, completed tasks, reverts, human fix time. One spreadsheet is enough without a dashboard. After two weeks you will see which work belongs to agents and which belongs to humans. The moment that split appears, AI use turns from hobby into engineering.
Further reading
- Judge AI model cost per successful task
- Why read pass@1, cost, and time together
- After coding agents comes control: Governed Software Factory
References
- GitLab Docs: Duo Agent Platform Impact dashboard
- GitLab: 19.4 brings new agentic automation at a lower cost
- GitLab/Forrester: 400% ROI with Duo Agent Platform
- GitLab: Duo Agent Platform docs
Go deeper with a course
Weighing cost against outcomes is ultimately a judgment problem: what to automate and what humans should review. Practice connecting recurring judgments to code, then designing confidence, policy, and human review together, continues directly from the four-number log above.