Cloudflare rebuilt Containers for agent sandboxes and the numbers moved a lot:
| Burst TTI, 100 concurrent sandboxes | Before | After | |
|---|---|---|---|
| Median | 4.049 s | 648 ms | 6.2× |
| p95 | 5.839 s | 910 ms | 6.4× |
| p99 | 6.717 s | 1129 ms | 5.9× |
Benchmark is ComputeSDK’s, run independently. Cloudflare separately reports starting 100,000 containers from one account in 5.387 seconds across six locations.
The engineering is unglamorous in the good way. Scheduling moved into the Durable Object so the global control plane is in the path less. Placement checks the local machine before widening its search. And the runtime stopped booting VMs from scratch — it restores a prepared one, reuses networking and filesystem setup, batches repeated work, and skips waiting on services the first command doesn’t need. There’s also a pre-distributed image, cloudflare/debian-trixie, carrying Debian Trixie Slim and Node 24.20.0 LTS, which takes the pull-and-unpack step out of the request path.
Worth noting what kind of improvement that is: no new isolation technology, no weaker boundary. They removed waiting.
Why I care about this number more than the usual benchmark
Cold start normally reads as a latency story. For agent sandboxes it’s a security story, and the mechanism is behavioural rather than technical.
At four seconds a sandbox, nobody creates one per task. They can’t — a multi-step agent run that spawns a sandbox per tool call would spend most of its wall time booting. So teams do the rational thing and amortise: keep a sandbox warm, reuse it across tasks, often across users, sometimes across tenants.
And now the properties you were buying are gone. Files written during task A are sitting there during task B. Environment variables, cached credentials, a poisoned node_modules from a package the agent installed two tasks ago — all still present. If one task gets compromised, the compromise persists in the reused sandbox instead of being thrown away with it.
That architecture didn’t come from a threat-model decision. It came from an invoice. The security posture was set by whatever made the latency budget work, and then the diagram was drawn afterwards to match.
This is the pattern I keep running into: an unaffordable control gets quietly dropped, and the dropping is invisible because nobody writes a design doc saying “we decided to share sandboxes across tenants.” At 648ms, the sandbox-per-task version is affordable, and the argument for sharing one has to be made on its merits instead of on arithmetic.
What it buys
A fresh sandbox per task. Whatever the agent did — files, installed packages, environment mutations — ends when the task does, by construction rather than by cleanup code. Cleanup code is where this usually fails; it’s always incomplete and nobody tests it.
Disposability for the genuinely risky step. Executing agent-generated code, running an untrusted tool, opening a document from outside — those are the operations worth a dedicated sandbox, and now you can give them one without a latency argument.
A sharper blast radius story. “This agent’s exposure ends when the task ends” is a sentence you can put in a review. “This agent runs in a long-lived sandbox we periodically recycle” is not the same sentence, and everyone in the room knows it.
A cheaper fan-out. Sub-second starts at high concurrency mean parallel sub-agents each get their own environment rather than sharing one and trampling each other’s working directory.
What it doesn’t buy
Starting a sandbox quickly says nothing about what the sandbox is allowed to do.
A disposable container holding a long-lived credential that can read your production database is a disposable process with access to your production database. The isolation boundary is now cheap; where you draw it is still entirely your problem, and that’s the part most teams have wrong. Egress policy, credential scope, and what data the thing can reach are untouched by any of these numbers.
Nor does a fresh sandbox protect you from the content the agent processes. Every page it loads and document it opens is untrusted text landing in a context window — prompt injection works identically in a one-second-old sandbox.
And there’s a trap in disposability itself: if the sandbox is thrown away, the forensic record goes with it. Discarding the environment is the point, but then whatever you needed to know about what happened in there has to have been streamed out while it ran. Cheap sandboxes without matching audit discipline means more incidents where the evidence was deleted as designed.
One operational note
The legacy Container and Sandbox classes are maintained only through 31 December 2026. Existing deployments keep running past that date; they just stop getting updates.
That’s the kind of line that’s easy to skim in a launch post and annoying to rediscover in February. If you have anything on the old classes, the migration is a known, dated piece of work — put it on the board now, while it’s cheap.
Related: Agent browsing got 3–7× cheaper and 1.8× slower · What a Container Actually Is: Namespaces, cgroups, and Layers · Agent sprawl is your next production incident
Comments