When an AI coding tool disappoints you, the first move is usually to switch models.
Luna disappoints → Sol
Sol disappoints → Astra
Astra disappoints → Claude
Sometimes swapping the model is the answer.
But when the same model performs very differently across projects, look outside the model.
A model never works alone
In agentic coding tools, the model is only one part of the real work.
The flow looks roughly like this.
User request
↓
Project instructions
↓
Model
↓
Tools
↓
Files / Terminal / Browser
↓
Tests / Checks
↓
Retry / Repair
The surrounding structure is broadly called the harness.
Even with a good model, results wobble when it:
- does not know the project rules
- does not know the test commands
- does not know which files are off-limits
- never verifies failures
- answers once and stops
Benchmarks now measure model plus agent together
Recent coding benchmarks often skip the model name alone.
Artificial Analysis runs its Coding Agent Index with each model inside a specific agent and harness — OpenAI models, for example, are measured together with the Codex harness. Terminal-Bench looks at how an agent behaves in a real terminal, not at plain code generation.
The reason is simple.
Model capability
≠
End-to-end agent performance
Build two environments with the same model
You can run a tiny experiment as a CodeBridge Mini Lab.
A. Minimal environment
Ask the AI only this.
Fix the login bug in this project.
B. Slightly reinforced harness
Give the same model these conditions.
Goal:
Fix the login validation bug.
Rules:
- Touch files outside src/auth only when needed
- Never change public API signatures
Verify:
- Run npm test
- Run npm run lint
- On failure, find the cause and fix once
Done means:
- All tests pass
- A new regression test is added
- Changed files and reasons are summarized
The model is identical.
What changed is the environment it works in.
What should you compare?
Run each condition three times and record this.
Tests passed
Retries
Files changed
Unexpected changes
Human correction needed
Time
If B is more stable, the gain came from harness improvement, not a model upgrade.
That does not mean B is always better. Too many rules can distract the model instead.
So the point is not "write more instructions." Find the failure point and add only the structure it needs.
The three things to add first
You can start with these three before any agent framework.
1. Project rules
Which directories are critical?
Which APIs must never change?
What is the code style?
2. Verification commands
npm test
pytest
cargo test
npm run lint
3. Done conditions
Tests pass
No needless file changes
No regressions in existing features
With just these three, AI shifts a little from "a tool that generates code" to "a tool that finishes work."
Why separate model upgrades from harness fixes
Because then failures are easier to diagnose.
Fixed by changing the model
→ possibly a capability problem
Fixed by adding a verification loop
→ possibly a process problem
Fixed by adding project rules
→ possibly a context problem
That separation also saves money.
With a solid harness, you may need the top-tier model on fewer tasks.
The direction OpenAI's Agents API shows
The Agents API, which OpenAI released on September 10, 2026, follows the same current.
OpenAI frames it as the managed Codex harness and infrastructure — session orchestration, context compaction, recovery, durable sessions, tool and MCP connections, and hosted sandboxes — delivered as an API.
The direction is shifting from serving one model call to serving the runtime where a model can work for a long time as an API.
Conclusion: before changing models, look at the failure structure
When AI coding struggles, you do not always need "a better model."
Try asking this once.
Is this model bad at the work, or did I give it a bad place to work?
The bigger the project grows, the more that question pays off.
And harness engineering starts less with a grand framework than with shrinking repeated failures through environment design.
Further reading
- What is harness engineering?
- SWE-bench vs Terminal-Bench vs ProgramBench
- OpenAI Agents API: the Codex harness as an API
References
- OpenAI: Introducing the Agents API
- Artificial Analysis: Coding Agent Index Methodology
- Anthropic: Effective harnesses for long-running agents
Go deeper with a course
If you want to practice turning repeated failures into rules, verification, and done conditions, a guided course builds the habit step by step.