GPT-6.1 Sol shipped with a quiet but important feature.

It is the multi-agent beta in the Responses API.

Instead of one model doing everything in order, a root agent can create subagents and split the work.

It looks simple in a diagram.

Single Agent

Task
  ↓
Agent
  ↓
Research
  ↓
Analyze
  ↓
Implement
  ↓
Review
  ↓
Done

With multi-agent, it becomes:

                    ┌→ Researcher
Task → Root Agent ──┼→ Implementer
                    └→ Reviewer
                          ↓
                    Merge results
                          ↓
                       Final

But this conclusion would be risky:

"Three agents must be three times faster."

In practice, that is often not true.

Multi-agent only helps certain kinds of work

The key question is can you split the work independently?

A pull request review splits well, for example.

Agent A
Accuracy / bugs

Agent B
Security

Agent C
Missing tests

All three see the same diff. They rarely need to wait for each other.

By contrast, this kind of work is hard to parallelize.

1. Design the API
2. Design the DB schema using the result of 1
3. Write the migration using the result of 2

Each step needs the previous result. Even with many subagents, you get waiting time.

So multi-agent works well when:

Parallelizable
+
Results can be merged later

OpenAI multi-agent is coordinated by a root

In the Responses API docs, the top-level agent is called /root.

When subagents are created, they can form a hierarchical path like this.

/root
├── /root/researcher
├── /root/reviewer
└── /root/reviewer/tester

A subagent can create its own subagents too.

So it is closer to an agent tree than to "three API calls."

The root agent breaks down the task, collects subagent results, resolves conflicts and duplicates, and then builds the final answer.

The Responses API defaults to 3 concurrent subagents

In the current multi-agent beta, the default for max_concurrent_subagents is 3.

An example looks like this.

from openai import OpenAI

client = OpenAI()

response = client.beta.responses.create(
    model="gpt-6.1-sol",
    input="""
    Review this PR from three angles.
    1. correctness
    2. security
    3. missing tests

    Merge duplicate comments,
    then organize by priority.
    """,
    multi_agent={
        "enabled": True,
        "max_concurrent_subagents": 3,
    },
    betas=["responses_multi_agent=v1"],
)

This is a simplified example for understanding the beta docs.

Because it is a beta API, check the latest SDK and docs before you adopt it.

More agents also means more cost

This is the easiest trap in parallel work.

With a single agent:

Context
→ analyze once
→ answer once

With multi-agent, the same context can enter several agents.

              ┌→ Context + reasoning A
Shared input ─┼→ Context + reasoning B
              └→ Context + reasoning C

                + Root synthesis

So wall-clock time can drop while token usage grows.

That is why you need at least four metrics.

Quality
Time
Tokens
Cost

If you only look at "it was faster," you see half the picture.

Duplicated work is also a cost

If you do not split roles clearly, this happens.

Agent A: research the whole repo
Agent B: research the whole repo
Agent C: research the whole repo

You parallelized the work but did the same work three times.

A good split looks like this:

A → auth module
B → billing module
C → tests / integration

Give each agent a non-overlapping scope, or:

A → correctness
B → security
C → test coverage

Give each agent a different view of the same target.

That is why task decomposition matters so much.

CodeBridge Mini Lab: single vs 3-agent PR review

This is the simplest comparison experiment.

Prepare one real PR diff.

A. Single agent

Review this PR and
find bugs, security issues, and missing tests.

B. Multi-agent

Review it in three roles.

1. correctness
2. security
3. missing tests

Remove duplicates across results and
organize by severity.

Then repeat each mode three times.

mode,run,valid_findings,false_positives,time_sec,input_tokens,output_tokens,cost
single,1,6,2,80,0,0,0
single,2,5,1,73,0,0,0
multi,1,8,2,55,0,0,0
multi,2,7,3,49,0,0,0

Ask these questions:

1. Did it find more valid issues?
2. Did false positives grow?
3. Did wall-clock time drop?
4. How much did total tokens and cost grow?
5. Were results consistent across runs?

If multi-agent is better, the reason should not be "more agents." It should be my task split well.

All subagents share the same tools

In Responses API multi-agent, the root and subagents can access the same tool set.

That is convenient, but it raises a permission question.

Does the reviewer need write tools?
Does the researcher need shell write access?

Give each role a clear scope. For important side effects, add a separate permission boundary in your application.

Multi-agent is a task-splitting feature. It does not design least privilege for you.

Each agent also manages context separately

When multi-agent is on, server-side compaction applies independently to the root and to each subagent context.

That matters in long tasks.

/root context

/root/researcher context

/root/reviewer context

They do not share one infinite history. Each agent keeps its own working context.

The current beta also has limits.

For example, with multi-agent:

  • reasoning.summary is not supported
  • max_tool_calls is not supported
  • direct use of the /responses/compact endpoint is not supported

Check constraints like these before you put a beta feature into production architecture.

Multi-agent naturally leads to graph engineering

When agents multiply, you soon ask:

Who runs first?
Who passes results to whom?
Where do you return on failure?
Who decides when two opinions conflict?

At that point it is not a prompt problem. It is a graph problem.

Planner
  ├─ Researcher A
  ├─ Researcher B
  └─ Implementer
          ↓
       Reviewer
       ↙      ↘
    pass      fail
     ↓         ↓
    end    Implementer

In multi-agent work, the key skill is not calling more models. It is designing this flow.

The bottom line: look at task decomposition before subagent count

The GPT-6.1 Sol multi-agent beta makes agent delegation much easier at the API level.

But the formula for good results is not this:

1 agent
→ 3 agents
→ 10 agents
→ better

This is closer:

Good task decomposition
+
Right amount of parallelism
+
Clear roles
+
Deduplication
+
Final verification
=
Useful multi-agent

At first, do not add agents. Check whether you can split one task into two or three independent pieces.

If you cannot, a single agent can be cheaper, simpler, and more stable.

Further reading

References

Go deeper with a course

If you want to practice task splitting, merging, and failure routing in real agent systems, start with a structured harness course.