Agent Sandboxes Start 6× Faster Now. Isolation Just Got Cheaper Than the Workaround.

Cloudflare cut median time-to-interactive from 4.0s to 648ms. The number matters because cold start is why teams share sandboxes they shouldn't share.


Sandbox cold start: at four seconds per sandbox, teams amortise the cost by reusing one long-lived sandbox across many tasks, so state and compromise carry over; at 648ms, a fresh sandbox per task becomes affordable and each task's blast radius ends when it does

Cloudflare rebuilt Containers for agent sandboxes and the numbers moved a lot:

Burst TTI, 100 concurrent sandboxesBeforeAfter
Median4.049 s648 ms6.2×
p955.839 s910 ms6.4×
p996.717 s1129 ms5.9×

Benchmark is ComputeSDK’s, run independently. Cloudflare separately reports starting 100,000 containers from one account in 5.387 seconds across six locations.

The engineering is unglamorous in the good way. Scheduling moved into the Durable Object so the global control plane is in the path less. Placement checks the local machine before widening its search. And the runtime stopped booting VMs from scratch — it restores a prepared one, reuses networking and filesystem setup, batches repeated work, and skips waiting on services the first command doesn’t need. There’s also a pre-distributed image, cloudflare/debian-trixie, carrying Debian Trixie Slim and Node 24.20.0 LTS, which takes the pull-and-unpack step out of the request path.

Worth noting what kind of improvement that is: no new isolation technology, no weaker boundary. They removed waiting.

Why I care about this number more than the usual benchmark

Cold start normally reads as a latency story. For agent sandboxes it’s a security story, and the mechanism is behavioural rather than technical.

At four seconds a sandbox, nobody creates one per task. They can’t — a multi-step agent run that spawns a sandbox per tool call would spend most of its wall time booting. So teams do the rational thing and amortise: keep a sandbox warm, reuse it across tasks, often across users, sometimes across tenants.

And now the properties you were buying are gone. Files written during task A are sitting there during task B. Environment variables, cached credentials, a poisoned node_modules from a package the agent installed two tasks ago — all still present. If one task gets compromised, the compromise persists in the reused sandbox instead of being thrown away with it.

That architecture didn’t come from a threat-model decision. It came from an invoice. The security posture was set by whatever made the latency budget work, and then the diagram was drawn afterwards to match.

This is the pattern I keep running into: an unaffordable control gets quietly dropped, and the dropping is invisible because nobody writes a design doc saying “we decided to share sandboxes across tenants.” At 648ms, the sandbox-per-task version is affordable, and the argument for sharing one has to be made on its merits instead of on arithmetic.

What it buys

A fresh sandbox per task. Whatever the agent did — files, installed packages, environment mutations — ends when the task does, by construction rather than by cleanup code. Cleanup code is where this usually fails; it’s always incomplete and nobody tests it.

Disposability for the genuinely risky step. Executing agent-generated code, running an untrusted tool, opening a document from outside — those are the operations worth a dedicated sandbox, and now you can give them one without a latency argument.

A sharper blast radius story. “This agent’s exposure ends when the task ends” is a sentence you can put in a review. “This agent runs in a long-lived sandbox we periodically recycle” is not the same sentence, and everyone in the room knows it.

A cheaper fan-out. Sub-second starts at high concurrency mean parallel sub-agents each get their own environment rather than sharing one and trampling each other’s working directory.

What it doesn’t buy

Starting a sandbox quickly says nothing about what the sandbox is allowed to do.

A disposable container holding a long-lived credential that can read your production database is a disposable process with access to your production database. The isolation boundary is now cheap; where you draw it is still entirely your problem, and that’s the part most teams have wrong. Egress policy, credential scope, and what data the thing can reach are untouched by any of these numbers.

Nor does a fresh sandbox protect you from the content the agent processes. Every page it loads and document it opens is untrusted text landing in a context window — prompt injection works identically in a one-second-old sandbox.

And there’s a trap in disposability itself: if the sandbox is thrown away, the forensic record goes with it. Discarding the environment is the point, but then whatever you needed to know about what happened in there has to have been streamed out while it ran. Cheap sandboxes without matching audit discipline means more incidents where the evidence was deleted as designed.

One operational note

The legacy Container and Sandbox classes are maintained only through 31 December 2026. Existing deployments keep running past that date; they just stop getting updates.

That’s the kind of line that’s easy to skim in a launch post and annoying to rediscover in February. If you have anything on the old classes, the migration is a known, dated piece of work — put it on the board now, while it’s cheap.


Related: Agent browsing got 3–7× cheaper and 1.8× slower · What a Container Actually Is: Namespaces, cgroups, and Layers · Agent sprawl is your next production incident

Frequently asked questions

How much faster are the rebuilt Cloudflare Containers?

On ComputeSDK's independent Burst TTI benchmark with 100 concurrent sandboxes, median time-to-interactive fell from 4.049 seconds to 648 milliseconds, a 6.2x improvement. The 95th percentile went from 5.839 seconds to 910 milliseconds and the 99th from 6.717 seconds to 1129 milliseconds. Cloudflare also reports starting 100,000 containers from a single account in 5.387 seconds across six locations.

What made them faster?

Three changes. Scheduling moved into the Durable Object so the global control plane is involved less. Placement now looks for capacity on the same machine first before widening the search within a location. And the runtime restores a prepared virtual machine instead of booting one from scratch, reusing networking and filesystem setup, batching repeated operations, and no longer waiting on services the first command does not need. A pre-distributed image, cloudflare/debian-trixie, also removes the download-and-unpack step from the request path.

Why does sandbox cold start affect security rather than just latency?

Because when a fresh sandbox costs four seconds, teams stop creating fresh ones. They keep a long-lived sandbox and reuse it across tasks, users and tenants, which means state from one task is visible to the next and a compromise persists instead of being discarded. The latency number sets the price of isolation, and when isolation is expensive people quietly buy less of it than their threat model assumes.

What does a faster sandbox not fix?

It does not change what the agent is permitted to reach. Network egress, credentials, and the data the sandbox is allowed to read are all unchanged by how quickly the sandbox starts. A disposable sandbox with a long-lived credential that can read your production database is still a disposable process with access to your production database. Cheap isolation makes per-task boundaries affordable; it does not decide where those boundaries go.

Comments