On October 7, Microsoft announced general availability for Microsoft Execution Containers (MXC). The same day, GitHub announced GA for the MXC-based local sandbox in Copilot.

The key sentence first:

An agent cannot be its own security authority. It must run within a boundary defined by the developer or organization and enforced independently of the agent itself.

In the era when agents execute code directly, dividing one file and one network line by policy has become a product feature. This post covers what MXC divides, how it switches on in Copilot, and how to carry it into your project.

Why it matters: a fast shortcut becomes an outage

The official blog's example is precise enough to reuse. Suppose a coding agent is asked to update a website.

Needed:
 - read and write the website repo
 - run build and test tools
 - read production server config (to understand deployment)

Never granted:
 - modifying production server config
 - touching personal files outside the repo
 - arbitrary network connections

Without a boundary, an agent goes wrong with good intentions. Editing the server config can look like the fastest fix, so it runs. Reasonable from the agent's view, beyond the authority the developer granted. The result is a production outage.

MXC's answer is plain. Read and write the repo, read the server config, grant nothing else. If the agent tries to modify the config, the container blocks it no matter what the model, generated code, plugin, or tool says. Authority is set outside the agent.

What MXC is: policy as JSON, enforcement by the OS

MXC is a policy-driven execution layer for untrusted code or dynamically generated workloads. In agent scenarios it can contain model output, plugins, tools, the harness, or the whole agent.

Developer declares (JSON policy + SDK):
  "this workload uses these files and these networks"

MXC enforces (at runtime):
  Windows → AppContainer
  macOS   → Seatbelt
  Linux   → Bubblewrap (mapped per backend)

Policy lives outside the workload
→ the agent cannot grant itself more access

Developers write one unified JSON schema. MXC maps platform isolation details to backends. The same containment model runs from local devices to the cloud, and Windows 365 support joined GA.

Pick your isolation: a four-step spectrum

Workloads need different isolation. A repo-coding agent wants low latency; an agent handling sensitive data wants a harder boundary.

Backend Availability Best suited for Characteristics
Process container Windows 11, macOS, Linux Lightweight containment for generated code and tools AppContainer / Seatbelt / Bubblewrap
Session container Windows 11 only Long-running agents needing a desktop Separate Windows account and session, isolated desktop, clipboard, UI, input
WSL container Windows 11 only Linux-first toolchains Linux execution via WSL
MicroVM Windows 11 and Linux, experimental Higher-risk workloads Hardware-backed virtualized boundary

Session containers are Windows-only and MicroVM is still experimental. These are not interchangeable; they form a spectrum picked by workload risk.

Five policy areas: files and networks are the core

Back to the website example. Five areas an MXC policy divides.

Area What it decides
Containment Which isolation runs the workload
Process Command, arguments, working directory, environment
File system Writable, read-only, and denied locations
Network Inbound and outbound, including loopback access
User interface Desktop and UI resource access

Copilot's defaults build intuition. With sandboxing on, the working directory is read-write and most of the rest is read-only or inaccessible. Shell commands and local MCP and language servers run inside the boundary by default. Built-in file tools are checked against policy inside the harness rather than OS-enforced, and remote MCP servers sit outside the local process sandbox with only connection policy checked. Inside and outside the boundary need separate thinking.

Combining with org policy: Intune is coming

MXC's good design decision is splitting developer declarations from organizational constraints.

Developer policy: "this agent workload needs these resources"
Org policy (Intune, coming soon): "our company never crosses this line"
→ the same agent runs inside different enterprise boundaries
→ developers stop hardcoding company posture into the app

The docs prescribe agent behavior when org policy blocks a resource. Never fail silently: explain the task could not finish inside the available permissions, request user or administrator action where supported, or pick a safe alternative. The no-silent-failure rule matches the MCP re-auth post.

Never enforce on day one: three modes

Least-privilege policy is hard to write blind. Nobody knows every resource an agent touches upfront. On Windows, process containers emit an activity report to help draft policy.

Mode Ungranted access Report Use
Enforcement Blocked None Run with the production policy
Learning Blocked and recorded Yes (JSON) Diagnose failures, verify least privilege
Permissive Allowed and recorded Yes Observe while authoring policy
Recommended order:
 Permissive first (collect what it touches)
   → Learning to tighten (check each denial is legitimate)
     → Enforcement for production (pin the policy)

First contact with MXC may block legitimately needed capabilities. Those denials show where the boundary needs widening. Granting only what is required is the design work itself.

Switching it on in Copilot: one /sandbox line

GitHub's announcement is short. Local sandboxing is GA in the Copilot CLI, the Copilot app, and VS Code sessions using Agent Host. MXC translates a common sandbox policy into native OS controls.

In the CLI:
  /sandbox → open settings
  turn on "Sandbox new sessions" → new sessions sandboxed by default

Defaults:
  working directory = read-write
  the rest = read-only or inaccessible

The ecosystem status is worth recording. Shipping with MXC now: Copilot, OpenAI Codex, OpenClaw, Replit, LM Studio, Unsloth AI, NVIDIA OpenShell, and others. Announced as coming: Claude Code, Box, Egnyte, Manus, Perplexity, Raycast, Simular, and more. Copilot users get it today; Claude Code users should file "coming soon" under roadmap, not reality. That separation is the basics of re-verification.

CodeBridge Mini Lab: build your sandbox

1. Pick 1 project (the repo your agent touches)
2. Write the policy in 5 lines:
   [ ] writable: repo paths (read-write)
   [ ] read-only: build and prod configs (read only)
   [ ] denied: personal docs and credentials (~/.ssh, token files)
   [ ] network: block inbound, allow outbound only where needed
   [ ] UI: deny desktop access (unless required)
3. Run once in Permissive → check the report for touched resources
4. Tighten in Learning → fold legitimate denials into policy
5. Pin in Enforcement → add 1 regression test
   (does the agent get blocked reaching forbidden zones)

It adds one sandbox line to the verification loop from using Claude Code on real projects. Sources writable, documents unreadable, credentials untouchable. That is the briefing's action.

Conclusion: the bigger the autonomy, the earlier the fence

One line to close.

Set the boundary before granting freedom. Blocked logs become the next policy.

MXC going GA is not one isolation technology. It signals agent execution moving into permission, identity, and management territory. Identity continues toward Entra and Agent 365, management toward Intune. Start small today. One sandbox that writes the work folder and blocks the rest. That fence is what lets you keep using agents long-term.

Further reading

References

Go deeper with a course

To practice controlling AI coding through verification and constraints, this course organizes project rules and permission environments with Claude Code, exactly like the sandbox design here.