GPT-6.1 Sol shipped with a quiet but important feature.
It is the multi-agent beta in the Responses API.
Instead of one model doing everything in order, a root agent can create subagents and split the work.
It looks simple in a diagram.
Single Agent
Task
↓
Agent
↓
Research
↓
Analyze
↓
Implement
↓
Review
↓
Done
With multi-agent, it becomes:
┌→ Researcher
Task → Root Agent ──┼→ Implementer
└→ Reviewer
↓
Merge results
↓
Final
But this conclusion would be risky:
"Three agents must be three times faster."
In practice, that is often not true.
Multi-agent only helps certain kinds of work
The key question is can you split the work independently?
A pull request review splits well, for example.
Agent A
Accuracy / bugs
Agent B
Security
Agent C
Missing tests
All three see the same diff. They rarely need to wait for each other.
By contrast, this kind of work is hard to parallelize.
1. Design the API
2. Design the DB schema using the result of 1
3. Write the migration using the result of 2
Each step needs the previous result. Even with many subagents, you get waiting time.
So multi-agent works well when:
Parallelizable
+
Results can be merged later
OpenAI multi-agent is coordinated by a root
In the Responses API docs, the top-level agent is called /root.
When subagents are created, they can form a hierarchical path like this.
/root
├── /root/researcher
├── /root/reviewer
└── /root/reviewer/tester
A subagent can create its own subagents too.
So it is closer to an agent tree than to "three API calls."
The root agent breaks down the task, collects subagent results, resolves conflicts and duplicates, and then builds the final answer.
The Responses API defaults to 3 concurrent subagents
In the current multi-agent beta, the default for max_concurrent_subagents is 3.
An example looks like this.
from openai import OpenAI
client = OpenAI()
response = client.beta.responses.create(
model="gpt-6.1-sol",
input="""
Review this PR from three angles.
1. correctness
2. security
3. missing tests
Merge duplicate comments,
then organize by priority.
""",
multi_agent={
"enabled": True,
"max_concurrent_subagents": 3,
},
betas=["responses_multi_agent=v1"],
)
This is a simplified example for understanding the beta docs.
Because it is a beta API, check the latest SDK and docs before you adopt it.
More agents also means more cost
This is the easiest trap in parallel work.
With a single agent:
Context
→ analyze once
→ answer once
With multi-agent, the same context can enter several agents.
┌→ Context + reasoning A
Shared input ─┼→ Context + reasoning B
└→ Context + reasoning C
+ Root synthesis
So wall-clock time can drop while token usage grows.
That is why you need at least four metrics.
Quality
Time
Tokens
Cost
If you only look at "it was faster," you see half the picture.
Duplicated work is also a cost
If you do not split roles clearly, this happens.
Agent A: research the whole repo
Agent B: research the whole repo
Agent C: research the whole repo
You parallelized the work but did the same work three times.
A good split looks like this:
A → auth module
B → billing module
C → tests / integration
Give each agent a non-overlapping scope, or:
A → correctness
B → security
C → test coverage
Give each agent a different view of the same target.
That is why task decomposition matters so much.
CodeBridge Mini Lab: single vs 3-agent PR review
This is the simplest comparison experiment.
Prepare one real PR diff.
A. Single agent
Review this PR and
find bugs, security issues, and missing tests.
B. Multi-agent
Review it in three roles.
1. correctness
2. security
3. missing tests
Remove duplicates across results and
organize by severity.
Then repeat each mode three times.
mode,run,valid_findings,false_positives,time_sec,input_tokens,output_tokens,cost
single,1,6,2,80,0,0,0
single,2,5,1,73,0,0,0
multi,1,8,2,55,0,0,0
multi,2,7,3,49,0,0,0
Ask these questions:
1. Did it find more valid issues?
2. Did false positives grow?
3. Did wall-clock time drop?
4. How much did total tokens and cost grow?
5. Were results consistent across runs?
If multi-agent is better, the reason should not be "more agents." It should be my task split well.
All subagents share the same tools
In Responses API multi-agent, the root and subagents can access the same tool set.
That is convenient, but it raises a permission question.
Does the reviewer need write tools?
Does the researcher need shell write access?
Give each role a clear scope. For important side effects, add a separate permission boundary in your application.
Multi-agent is a task-splitting feature. It does not design least privilege for you.
Each agent also manages context separately
When multi-agent is on, server-side compaction applies independently to the root and to each subagent context.
That matters in long tasks.
/root context
/root/researcher context
/root/reviewer context
They do not share one infinite history. Each agent keeps its own working context.
The current beta also has limits.
For example, with multi-agent:
reasoning.summaryis not supportedmax_tool_callsis not supported- direct use of the
/responses/compactendpoint is not supported
Check constraints like these before you put a beta feature into production architecture.
Multi-agent naturally leads to graph engineering
When agents multiply, you soon ask:
Who runs first?
Who passes results to whom?
Where do you return on failure?
Who decides when two opinions conflict?
At that point it is not a prompt problem. It is a graph problem.
Planner
├─ Researcher A
├─ Researcher B
└─ Implementer
↓
Reviewer
↙ ↘
pass fail
↓ ↓
end Implementer
In multi-agent work, the key skill is not calling more models. It is designing this flow.
The bottom line: look at task decomposition before subagent count
The GPT-6.1 Sol multi-agent beta makes agent delegation much easier at the API level.
But the formula for good results is not this:
1 agent
→ 3 agents
→ 10 agents
→ better
This is closer:
Good task decomposition
+
Right amount of parallelism
+
Clear roles
+
Deduplication
+
Final verification
=
Useful multi-agent
At first, do not add agents. Check whether you can split one task into two or three independent pieces.
If you cannot, a single agent can be cheaper, simpler, and more stable.
Further reading
- What is graph engineering?
- Qwen Code: coding agents started handing work to other agents
- Why the harness changes results more than the model
- Why you should read Pass@1, cost, and time together in coding agents
References
- OpenAI API: Responses Multi-agent
- OpenAI API Changelog — GPT-6.1 Sol Multi-agent beta
- OpenAI API: GPT-6.1 Sol
Go deeper with a course
If you want to practice task splitting, merging, and failure routing in real agent systems, start with a structured harness course.