Reinforcement learning does not run at average speed. The slowest rollout in a parallel step sets the step time. AWS proved that obvious statement with numbers in an October 9 Compute Blog report. Pack RL sandboxes tightly into Firecracker microVMs and the median (p50) stays flat while the tail (p99) explodes. More than 100,000 samples across 500+ configurations were measured, and p99 spiked the moment density passed 1.0 microVM per vCPU.
The rig: measure from the metal up
Start with background. RLVR-style training needs thousands of isolated sandboxes in parallel. Faster dispatch and faster recycling mean more steps per hour and lower cost per step. AWS published this reference shape:
EC2 metal host + host agent (DaemonSet)
- CPU topology discovery, one cgroup per microVM
- Firecracker jailer lifecycle, CPU pinning, NUMA placement
- warm pool of pre-booted microVMs for dispatch + telemetry
- a Kubernetes cluster above
Measurements ran on m8i, m8a, and m8g metal-48xl machines, with published numbers centered on m8i and m8a. The premises are explicit: Linux kernel 6.2+ (cgroup isolated mode), Firecracker 1.7+ with jailer, Go 1.22+ to build labsweep. Only CPU and memory placement were swept; storage (EBS vs local SSD) and the LLM-endpoint network were held constant. Each trial held one fixed density, so dynamic churn is out of scope.
The density knobs read clearly as a table:
Core pinning (cpuset.cpus) → dedicated physical core, kills scheduler tail
NUMA confinement (cpuset.mems)→ keep allocation local, check distance matrix
SMT on/off → m8i only (m8a/m8g have no SMT)
L3 cache domain pinning → matters past ~4 MiB working set per microVM
cgroup isolated partition → removes workqueue interference
Guest memory size → M-series baseline is 4 GiB per vCPU
Copied straight from the post, this example gives the feel:
mkdir -p /sys/fs/cgroup/microvm/vm-012
echo "12" > /sys/fs/cgroup/microvm/vm-012/cpuset.cpus
echo "0" > /sys/fs/cgroup/microvm/vm-012/cpuset.mems
firecracker-jailer --id vm-012 \
--exec-file /usr/bin/firecracker \
--parent-cgroup microvm/vm-012 \
--uid 1000 --gid 1000
The labsweep harness (a Go tool) has a fixed job: read topology from sysfs, plan a configs-by-densities matrix, launch jailed Firecrackers at each density with bounded concurrency, run the workload in every guest at once, and record per-microVM latency. It emits raw samples plus percentile aggregates, variance, and full provenance (instance, CPU, kernel, Firecracker version), and reports failed samples instead of hiding them. A typical m8a run looks like this:
labsweep -configs A,B \
-densities 90,120,150,186 \
-repeats 5 \
-resident-kib 131072 \
-steps 14 \
-write-kib 4096
Four workloads were swept: short verification (a test suite), build-heavy (compile and link), full 14-turn replayed RL episodes, and a synthetic compute loop as control. That split matters below.
The 1.0 wall: what happens past one per vCPU
The headline result is portable because it is a ratio, not an absolute count. The tail stays flat until 1.0 microVM per vCPU, then p99 spikes across the next ~15% of density onto a new plateau. The median does not move. Here is the appendix table (m8i, SMT off, 96 physical cores, unpinned config A):
Density Ratio p50 p99 Group amplification (G=8 / 16 / 64)
45 0.47x 7,768us 7,857us 1.01 / 1.01 / 1.01
90 0.94x 7,740us 8,035us 1.01 / 1.01 / 1.04
100 1.04x 7,740us 8,601us 1.01 / 1.03 / 1.13
110 1.15x 7,737us 9,909us 1.01 / 1.04 / 1.32
135 1.44x 7,743us 15,601us 1.03 / 1.20 / 2.01
186 1.94x 7,744us 15,697us 1.22 / 1.56 / 2.03
p50 sits near 7.7ms throughout, while p99 rises 7% at 1.04x, 23% at 1.15x, and roughly doubles from 1.44x. The amplification column on the right is the real trap. In GRPO-style training that waits for the slowest group member, bigger groups hit p99 events more often — at G=64, nearly 47% of steps meet one. Median dashboards hide this. Watch the amplification factor, as the model speed and latency guide argues for averages and tails together.
Pinning only pays on real workloads
The second result is a lesson about toys versus real sandboxes. On the synthetic compute loop, core pinning showed almost no difference at equal density. On real sandboxes, p99/p50 tightened toward ~1.1x, and the advantage grew with density.
SMT splits by workload too. For I/O-bound work (most RL sandboxes: files, process forks, git waits), SMT-on wins. On the same m8i host: verification 1.14x, build 1.13x, replayed RL episodes 1.17x throughput, stretching to 1.49x at high density because the vCPU count stays above the 1.0 line. For compute-bound loops, SMT-off gives a tighter, more predictable p99 spread, and at or below 1.0x it is slightly better per task. There is no universal optimum — measure at your target density.
Finally the memory wall. M-series machines carry 4 GiB per vCPU, so every tested workload exhausted vCPUs first. Needing more memory per vCPU means moving to a memory-optimized family, which cuts vCPUs per host and lowers maximum density. That is behind AWS's advice to tune from your own workload, not the spec sheet: SMT, cache, and NUMA differ per instance.
CodeBridge Mini Lab: measure p50, p99, and failures together
Turn the briefing's action item into an experiment order. Raise concurrent sandboxes while watching three things at once:
1. Convert density to a ratio:
- confirm host vCPUs (the denominator changes with SMT on/off)
- sweep 0.5x → 0.94x → 1.04x → 1.15x → 1.44x → 1.9x
2. Repeat 5+ times per density:
- p50, p99, p99/p50, group amplification max(group)/p50 (your G among 8/16/64)
- keep failed samples as their own count, never drop them
- hold EBS, LLM endpoint, and version provenance fixed and recorded
3. Decision lines:
- p50 flat but p99/p50 > 1.2 or G=64 amplification > 1.1 means past the wall
- step back one notch to the densest point where amplification stays flat
- measure with a real workload (synthetic loops hide pinning effects)
The warm pool is a separate knob. It governs dispatch latency, so keep it out of the density sweep. For cleanup, bundle cdk destroy, metal termination, and S3 rootfs plus CloudWatch log-group deletion as one set.
Conclusion: watch the slowest, not the average
One line to summarize.
Parallel training runs at tail speed, not average speed.
The 1.0-per-vCPU wall, pinning that fools synthetic loops, SMT that flips by workload, and a memory layout that exhausts vCPUs first all say the same thing. Oversubscription looks like headroom only until the tail owns training time. Convert today's concurrent sandbox count into a ratio. Past 1.0 without watching p99, one slow task is spending money on every step. Keep the agent sandbox security guide and the AI observability post nearby to turn this experiment into operating metrics.
Further reading
- AI model speed and latency: averages and tails together
- A security guide to AI agent sandboxes
- Measuring AI agents by cost per successful task
References
- AWS Compute Blog: Scaling RL rollouts with Firecracker on Amazon EC2 metal
- Firecracker Docs: Getting started
- AWS Docs: EC2 bare metal instance types
Go deeper with a course
Letting AI work inside isolated execution environments, controlled by permissions and checks, continues directly from this post's sandbox operations. To leave vibe-based AI coding for verified development flows, start here.