When an AI coding tool disappoints you, the first move is usually to switch models.

Luna disappoints → Sol
Sol disappoints → Astra
Astra disappoints → Claude

Sometimes swapping the model is the answer.

But when the same model performs very differently across projects, look outside the model.

A model never works alone

In agentic coding tools, the model is only one part of the real work.

The flow looks roughly like this.

User request
    ↓
Project instructions
    ↓
Model
    ↓
Tools
    ↓
Files / Terminal / Browser
    ↓
Tests / Checks
    ↓
Retry / Repair

The surrounding structure is broadly called the harness.

Even with a good model, results wobble when it:

  • does not know the project rules
  • does not know the test commands
  • does not know which files are off-limits
  • never verifies failures
  • answers once and stops

Benchmarks now measure model plus agent together

Recent coding benchmarks often skip the model name alone.

Artificial Analysis runs its Coding Agent Index with each model inside a specific agent and harness — OpenAI models, for example, are measured together with the Codex harness. Terminal-Bench looks at how an agent behaves in a real terminal, not at plain code generation.

The reason is simple.

Model capability
≠
End-to-end agent performance

Build two environments with the same model

You can run a tiny experiment as a CodeBridge Mini Lab.

A. Minimal environment

Ask the AI only this.

Fix the login bug in this project.

B. Slightly reinforced harness

Give the same model these conditions.

Goal:
Fix the login validation bug.

Rules:
- Touch files outside src/auth only when needed
- Never change public API signatures

Verify:
- Run npm test
- Run npm run lint
- On failure, find the cause and fix once

Done means:
- All tests pass
- A new regression test is added
- Changed files and reasons are summarized

The model is identical.

What changed is the environment it works in.

What should you compare?

Run each condition three times and record this.

Tests passed
Retries
Files changed
Unexpected changes
Human correction needed
Time

If B is more stable, the gain came from harness improvement, not a model upgrade.

That does not mean B is always better. Too many rules can distract the model instead.

So the point is not "write more instructions." Find the failure point and add only the structure it needs.

The three things to add first

You can start with these three before any agent framework.

1. Project rules

Which directories are critical?
Which APIs must never change?
What is the code style?

2. Verification commands

npm test
pytest
cargo test
npm run lint

3. Done conditions

Tests pass
No needless file changes
No regressions in existing features

With just these three, AI shifts a little from "a tool that generates code" to "a tool that finishes work."

Why separate model upgrades from harness fixes

Because then failures are easier to diagnose.

Fixed by changing the model
→ possibly a capability problem

Fixed by adding a verification loop
→ possibly a process problem

Fixed by adding project rules
→ possibly a context problem

That separation also saves money.

With a solid harness, you may need the top-tier model on fewer tasks.

The direction OpenAI's Agents API shows

The Agents API, which OpenAI released on September 10, 2026, follows the same current.

OpenAI frames it as the managed Codex harness and infrastructure — session orchestration, context compaction, recovery, durable sessions, tool and MCP connections, and hosted sandboxes — delivered as an API.

The direction is shifting from serving one model call to serving the runtime where a model can work for a long time as an API.

Conclusion: before changing models, look at the failure structure

When AI coding struggles, you do not always need "a better model."

Try asking this once.

Is this model bad at the work, or did I give it a bad place to work?

The bigger the project grows, the more that question pays off.

And harness engineering starts less with a grand framework than with shrinking repeated failures through environment design.

Further reading

References

Go deeper with a course

If you want to practice turning repeated failures into rules, verification, and done conditions, a guided course builds the habit step by step.