The agent thinks; the sandbox executes. This piece breaks the 2026 sandbox market into four tracks—specialized hosted platforms, cloud-vendor offerings, local agent isolation, and open-source self-hosting—explains what really separates containers, gVisor, Kata, and microVMs, why hourly rates are not comparable across vendors, how well-known agent products actually build this layer, and ends with selection and security checklists you can use directly.
Conclusions First
- A sandbox is not an agent framework; it is the agent's "execution computer." The agent handles thinking and decisions; the sandbox confines code, shell, files, processes, and network inside a controlled environment.
- By 2026 the market has settled into four competing tracks:
- Specialized hosted platforms: E2B, Daytona, Runloop, Blaxel, Fly.io Sprites, CodeSandbox SDK;
- Cloud vendors and developer platforms: Modal, Vercel Sandbox, Cloudflare Sandbox SDK, AWS AgentCore, GKE/Cloud Run Sandboxes, Azure Dynamic Sessions, CoreWeave;
- Local agent isolation: Docker Sandboxes, Anthropic Sandbox Runtime, plus the OS sandboxes shipped inside Codex, Claude Code, and DeepSeek Harness;
- Open-source self-hosting: E2B, OpenSandbox, Tencent CubeSandbox, Kimi AgentENV, BoxLite, NVIDIA OpenShell, Kubernetes Agent Sandbox, Microsandbox, AIO Sandbox—composed with Firecracker (started and open-sourced by AWS), gVisor (started and open-sourced by Google), or Kata Containers (a community project hosted by the OpenInfra Foundation).
- For a new team, E2B is the safest general-purpose starting point. If you need a long-lived "work computer," look first at Sprites, Blaxel, Daytona, and Runloop. Inside an existing cloud ecosystem, prefer that ecosystem's native offering. If you need GPUs, start with Modal, Northflank, or Daytona.
- Do not treat an ordinary Docker container as a security sandbox. Containers share the host kernel. Facing public users or LLM-generated untrusted code, add at least gVisor (Google) or Kata Containers (OpenInfra Foundation community); for higher risk, use a microVM such as Firecracker (AWS), or a lightweight VM with Hyper-V (Microsoft) isolation.
- You must separate agent harness from sandbox. Codex, Claude Code, Pi, and DeepSeek Harness are harnesses responsible for the loop, the context, and tool calls; E2B, Sprites, CubeSandbox, and the like are the execution isolation layer. One product can ship a lightweight local sandbox and also route tool execution to a remote microVM.
- Real security is more than "no escapes": you also have to solve network exfiltration, secret leakage, and resource abuse. A microVM mainly addresses the first of the three.
1. Core Concepts and Isolation Mechanisms
A common architecture looks like this:
User request
↓
Agent Runtime (model calls, planning, tool selection, session state)
↓ Sandbox API
Control plane (create, pause, snapshot, destroy, quota, audit)
↓
Isolated execution environment (shell / files / processes / browser / desktop)
├─ Filesystem and snapshots
├─ Network egress policy
├─ Secret proxy / short-lived credentials
└─ CPU, memory, disk, timeout limits
↓
Isolation substrate (container / gVisor / Kata / microVM)A sandbox has to answer at least six questions:
- Isolation boundary: a process, a container, a user-space kernel, or an independent microVM kernel?
- State model: destroyed after one use, or able to pause, resume, snapshot, and fork?
- Network model: full internet access by default, or deny-by-default with per-domain/CIDR allowlisting?
- Secret model: do real keys ever enter the sandbox? Can a proxy inject them only as a request leaves the sandbox?
- Resource model: are CPU, memory, disk, process count, and maximum runtime bounded?
- Return model: does the agent deliver stdout, files, a patch, a Git branch, a web preview, or a full snapshot?
The point that gets confused most often: the harness decides "what to do," the sandbox constrains "what can be done and where"; containers, gVisor, Kata, and microVMs are the isolation substrates a sandbox can choose from.
1.1 What Actually Separates the Isolation Mechanisms
| Term | Vendor / primary maintainer | In plain terms | Shares host kernel | Typical order of isolation strength | Startup / compatibility | Common examples |
|---|---|---|---|---|---|---|
| Process restriction | No single vendor; seccomp, namespaces, and Landlock are Linux kernel mechanisms, chroot comes from Unix | Limits which files a process can see and which system capabilities it can call | Yes | Low to medium | Extremely fast, but the boundary depends on OS configuration | seccomp, chroot, namespaces, Landlock |
| OS sandbox | Shipped with the OS or by the community: Seatbelt (Apple), bubblewrap (community), restricted token (Microsoft) | Uses native OS security mechanisms to draw a permission boundary around a process | Yes | Medium; suitable for protecting a personal dev machine | Very fast, native feel | macOS Seatbelt, Linux bubblewrap, Windows restricted token |
| Container | Generic architecture; Docker from Docker; containerd started by Docker and Kubernetes by Google, both now CNCF projects | Gives a process its own file, network, and process view while sharing the host kernel | Yes | Medium; potentially lower when misconfigured | Fast, best ecosystem | Docker, containerd, Kubernetes Pod |
| gVisor | Started and open-sourced by Google, maintained by Google and the community | Inserts a user-space kernel between the application and the host kernel, intercepting and reimplementing a large share of syscalls | Partly still depends on the host, but reduces the direct attack surface | Medium to high | Some compatibility and I/O cost beyond a container | The usual substrate under GKE Agent Sandbox |
| Kata Containers | An independent open-source community hosted by the OpenInfra Foundation | Looks like a container, actually puts each Pod/container inside a lightweight VM | No | High | Keeps OCI/K8s workflows, but costs more than a plain container | CoreWeave Sandboxes, enterprise K8s |
| microVM | Generic architecture, no single vendor; Firecracker started by AWS | A stripped-down virtual machine; every sandbox gets its own guest kernel | No | High | Better suited than a traditional VM to fast, high-volume creation | Firecracker (AWS), CubeSandbox (Tencent Cloud), Sprites (Fly.io) |
| Traditional VM | Generic architecture, no single vendor; Hyper-V from Microsoft, VMware products now under Broadcom | A complete virtual computer with richer device and management capabilities | No | High | Usually heavier, but compatible with Windows, complex drivers, and legacy ops | Cloud instances, Hyper-V/VMware VMs |
"Strength" here is only the usual architectural order of magnitude, not a security certification. A "strongly isolated" platform with a wide-open network, bare secrets, and no patching can easily be less secure overall than a carefully configured lightweight option.
Unfamiliar terms are defined in Appendix A.
2. The Current Landscape
2.1 Specialized Hosted and General-Purpose Platforms
| Option | Core positioning | Isolation and state | Public pricing (summary) | Open source and GitHub stars | Public users / adoption | Best suited for |
|---|---|---|---|---|---|---|
| E2B | General-purpose AI agent cloud sandbox; one of the de facto standards for code interpreters | Firecracker microVM; templates, pause/resume, persistent snapshots; up to 1h (Hobby) / 24h (Pro) | Hobby $0 + usage, one-time $100 credits; Pro $150/mo + usage. CPU $0.000014/vCPU·s, memory $0.0000045/GiB·s | Full infrastructure is self-hostable, Apache-2.0; e2b-dev/E2B 13,625★ | Manus (self-hosted), Hugging Face Open-R1, Groq Compound | General code execution, data analysis, RL/eval, high concurrency, teams wanting a self-hosting escape hatch |
| Daytona | A fast, composable agent computer; Linux/Windows/container/GPU in one | Linux containers by default, plus dedicated Linux/Windows VMs and GPUs; containers claim <90ms; snapshot, fork, pause | Pay-as-you-go with no base subscription; $0.0504/vCPU·h, $0.0162/GiB RAM·h, $0.000108/GiB disk·h; up to $200 credits | The current production platform is closed source. The historical repository is still public but no longer maintained; daytonaio/daytona 71,860★. The company announced the move to closed source in June 2026 | Cognition Devin Outposts, LangChain Open SWE, Brainbase Universal Harness | Low-latency interaction, persistent dev environments, Windows or GPU, computer use; do not mistake the old repository for the current product |
| Runloop | A devbox platform built for coding agents, benchmarking, and training/eval | Dedicated VMs; blueprints, disk snapshots, suspend/resume, network policy, secret/LLM gateway | Basic $0 + usage; Pro $250/mo + usage. $0.108/CPU·h, $0.0252/GB RAM·h | Platform is closed source, only SDK/API clients are public; no comparable platform star count | Accrual, ION, Trajectory; a published Trajectory case burst to 10,000+ concurrent devboxes | Coding agents, SWE-bench/private benchmarks, parallel attempts, enterprise credential proxying |
| Blaxel | A long-lived, idle-to-zero autonomous agent runtime | One microVM per agent; state persists continuously; claims ~25ms resume; network gateway and secret proxy are first-class | No base fee, up to $200 credits. Sandbox $0.0000115/GB RAM·s (active), snapshots $0.20/GB·mo | Platform is closed source; a public SDK and examples do not make the platform open source; platform star count not applicable | Webflow, CodSpeed, Runwork | Agents that wait a long time on external events and need cheap idle-to-zero plus network governance |
| Fly.io Sprites | A persistent, complete Linux computer for coding agents | Hardware-isolated microVM; persistent ext4, unlimited checkpoints, auto-sleep, dedicated URL | CPU $0.07/CPU·h, memory $0.04375/GB·h, billed on actually active resources; hot/cold storage billed separately | Service is closed source; sprites-js is only an MIT SDK, 34★ | SpriteDoc, Nous Research Hermes | Long-running coding tasks, webhooks/background services, moving an agent off a personal machine onto a persistent remote computer |
| CodeSandbox SDK | Programmable VMs grown out of an online IDE, emphasizing fork, hibernate, and collaboration | microVM; snapshot, clone, hibernate; resume typically within seconds | Build is free with 40 VM hours/mo and 10 concurrent; Scale $170/mo with 160 hours and 250 concurrent; extra resources about $0.0446/vCPU·h + $0.0149/GB·h | The SDK service/VM platform is closed source; stars on the old editor repository do not represent the SDK | Together AI, Superblocks, HeroUI | Online IDEs, education, code playgrounds, multiplayer development; a relatively high entry price for production |
| Northflank Sandboxes | A complete application platform: sandbox + PaaS + databases + GPU + BYOC | microVM-backed containers with persistent volumes, GPUs, BYOC, multi-cloud | Development sandboxes are free with limits; PAYG $0.01667/vCPU·h + $0.00833/GB·h, disk $0.15/GB·mo, egress $0.06/GB | Platform is closed source; no comparable platform star count | cto.new | Teams that need to run not just code but also APIs, databases, GPUs, and a private cloud/VPC alongside it |
| Modal Sandboxes | Sandboxes on a serverless Python/ML compute platform, strong on GPU and batch concurrency | Container sandboxes plus VM sandboxes; images, volumes, notebooks, GPUs; billed on max(request, actual) | Starter is $0 and includes $30 of compute per month; sandbox $0.00003942/physical core·s, $0.00000667/GiB·s, GPUs billed separately | Platform is closed source; modal-client is a client-only Apache-2.0 library, 513★ | Lovable, Ramp Inspect, Cognition | Data science, GPUs, batch processing, RL environments, teams already on Modal |
| Vercel Sandbox | Untrusted code execution and previews, native to Vercel/Next.js | Firecracker microVM; OCI images, root/sudo, Docker-in-Docker, snapshots; up to 24h on Pro | Pro $20/mo; active CPU $0.128/vCPU·h, memory $0.0212/GB·h, egress $0.15/GB; CPU wait time is not billed | Platform is closed source; vercel/sandbox is an Apache-2.0 SDK/CLI, 194★ | Notion Workers, Xata, Okara | AI UI builders on Vercel, code previews, agents that spend most of their time waiting on I/O |
| Cloudflare Sandboxes | Workers-native edge code execution, built from a Worker + Durable Object + Container | Every sandbox container runs in its own VM; auto-sleep, file/process/port APIs; the stable release is available while the 1.0 API is in a preview migration period | Workers Paid from $5/mo; $0.000020/vCPU·s (active), $0.0000025/GiB·s, $0.00000007/GB disk·s | The hosted platform is closed source; sandbox-sdk is an Apache-2.0 SDK, 1,121★ | — | Workers/Cloudflare-native applications, edge distribution, short and bursty execution |
| Docker Sandboxes | Puts an entire local agent—Claude Code, Codex, Gemini CLI, OpenCode—inside a dedicated microVM | One microVM per agent with its own Linux kernel and Docker daemon; the workspace can be mounted directly or privately cloned; a host proxy injects credentials | The sbx CLI is free for both personal and commercial use; org-level policy and audit require contacting sales |
Not an open-source implementation; docker/sbx-releases only publishes binaries and tracks issues, 344★ | — | Individuals and teams running coding agents safely on a local machine, especially when the agent also builds containers |
| Railway Sandboxes | Short-lived Linux work environments inside the Railway application platform; the default image ships Claude Code, Codex, OpenCode, and Pi | Templates, checkpoints, forks, long commands that survive disconnects, port forwarding, project-private networking; currently in Priority Boarding/beta | Billed as Railway VM usage: CPU and RAM both $0.001157/unit·min, egress $0.05/GB; billed only while the VM runs | Service is closed source, the TypeScript SDK is open; platform star count not applicable | — | Agents that need a temporary computer while also reaching databases and internal services in the same Railway project |
2.2 Cloud Vendors and Developer Platforms
| Option | Positioning and characteristics | Pricing | Open source / stars | Public users / adoption | Use cases |
|---|---|---|---|---|---|
| AWS Bedrock AgentCore Code Interpreter | AWS-native managed code interpreter; Python/JS/TS preinstalled, CloudTrail, enterprise IAM; CPU billed only on active consumption | $0.0895/vCPU·h + $0.00945/GB·h, per second, no minimum commitment; network billed as EC2 | Service is closed source, — | Swisscom and Iberdrola use AgentCore; not necessarily the Code Interpreter specifically | AWS enterprises, data analysis agents, IAM/audit/regional compliance requirements—but not arbitrary full Linux workstations |
| GKE Agent Sandbox | Kubernetes-native CRDs, warm pools, snapshots; isolates Pods with gVisor; more of a platform-engineering building block | Agent Sandbox carries no extra charge; you pay for GKE nodes, storage, and network | The GKE managed capability is closed source; upstream kubernetes-sigs/agent-sandbox is Apache-2.0, 3,702★ | Lovable, LangChain | Existing GKE/Kubernetes platform teams that need control over scheduling, storage, and network policy plus large warm pools |
| Azure Container Apps Dynamic Sessions | Hyper-V-isolated pre-warmed session pools; either a built-in interpreter or a custom container | Varies by region and contract; the built-in interpreter is billed per session-hour, rounded up to a full hour per allocation; custom containers use the Dedicated plan | Service is closed source, — | Microsoft Copilot Code Interpreter; the announcement disclosed 400,000+ sessions per day already in 2024 | Azure/Semantic Kernel/LangChain enterprises that want pre-warming and Hyper-V isolation, provided they budget carefully for the one-hour billing granularity |
| Cloud Run Sandboxes | Creates a nested sandbox inside the same Cloud Run instance hosting the agent, avoiding a separate cloud instance per execution; layered isolation that by default hides the parent workload, environment variables, secrets, and metadata; supports background sandboxes, persistent directories, and tar snapshots; public preview | No separate surcharge found; billed as Cloud Run CPU, memory, and network | Service is closed source, — | — | Agents already running on Cloud Run that want sub-second local creation and minimal IAM exposure |
| CoreWeave Sandboxes | A managed execution layer for RL rollouts, agent harnesses, and large-scale evaluation; can land on your own CKS or be used serverless through W&B; the CKS mode creates Pods in the customer cluster under namespace, network, and resource policy, while the serverless product page states Kata VM isolation | The CKS mode is in preview for existing customers and reuses purchased capacity; the serverless mode is offered through W&B; no simple retail pricing published | Service is closed source, — | — | Teams already training models on CoreWeave that want rollout/eval close to their models, data, and existing compute |
On adoption: at most three vendor- or customer-disclosed reference cases are listed per option; "—" only means there is not yet a sufficiently clear independent production disclosure. "Uses AgentCore" also does not necessarily mean using its Code Interpreter.
3. Why You Cannot Compare on "Per Hour" Alone
Vendors use at least four different billing models:
- Billed on provisioned resources for as long as the sandbox is alive: E2B, Daytona, Runloop, CodeSandbox, Northflank, Railway. While the agent waits on the model or the network, you usually still pay for memory and CPU.
- CPU billed on active usage only, memory on wall clock or peak: Vercel, Cloudflare, AWS AgentCore. Usually cheaper for I/O-heavy agents.
- Billed on actual/active resources with automatic idle-to-zero: Sprites, Blaxel; suited to long-lived agents with a low duty cycle.
- Billed on max(requested, actual) resources: Modal. Setting the request too high wastes money outright.
Hourly rates under different models cannot be ranked directly. For budgeting, use:
Total cost = base plan
+ creation/build charges
+ active CPU
+ wall-clock memory
+ hot/cold disk and snapshots
+ image storage
+ network egress and static IP/gateway
+ concurrency scaling charges
+ ancillary resources: logs, Durable Objects, object storage4. Open-Source Self-Hosted Options
4.1 Self-Hosting Comparison
Open source and self-hostable can mean a complete multi-tenant platform or merely an embeddable single-machine runtime. The table below separates deployment shape, entry point, and complexity to make them comparable. "Deployment complexity" is a relative assessment based on prerequisites, component count, and production operational requirements—not an official rating from any project.
| Option (vendor / primary maintainer) | Positioning and isolation | Shape | Deployment entry point / main prerequisites | Deployment complexity | License / stars | Who it suits |
|---|---|---|---|---|---|---|
| E2B The E2B team and community |
The same sandbox API, SDKs, control plane, and self-hosted infrastructure as the hosted service; Firecracker, supports AWS/GCP | Complete platform | Self-hosting guide; needs a cloud account, Cloudflare/domain, PostgreSQL, Packer, Terraform, Docker; AWS compute nodes require nested virtualization | High | Apache-2.0; 13,625★ | Teams that want SaaS first and BYOC/self-hosting later, and can live with Terraform/Nomad/virtualization operations |
| OpenSandbox Started by Alibaba, now maintained by the opensandbox-group community |
A general protocol, multi-language SDKs, lifecycle control plane, and Docker/K8s runtimes; composable with gVisor, Kata, Firecracker | Single machine or Kubernetes cluster | Single-machine/Docker configuration, Kubernetes Helm; needs Docker, the cluster path needs K8s/Helm, strong isolation needs a separate runtime | Low to high Low on Docker; medium to high with K8s + strong isolation |
Apache-2.0; 14,884★ | Platform teams avoiding vendor API lock-in who need multiple scenarios and runtimes |
| TencentCloud CubeSandbox Tencent Cloud |
An E2B-SDK-compatible single-machine/cluster platform with control plane, networking, and snapshots; RustVMM + KVM microVMs | Single machine or cluster | Quick start (Chinese); needs Linux/root, glibc 2.31+, KVM or PVM, and preferably a prepared XFS data disk | Medium to high Medium single-machine; high for cluster and public-network governance |
Apache-2.0; 11,571★ | Self-hosting teams wanting a domestic Chinese option, hardware isolation, E2B API compatibility, and a complete cluster architecture |
| NVIDIA OpenShell NVIDIA and the community |
A security policy and credential governance layer; optional Docker, Podman, K8s, or VM drivers | Local or cluster policy runtime | Installation, Quickstart; choose the Docker, Podman, K8s, or microVM driver | Low to medium Low for local containers; medium for microVM/K8s |
Apache-2.0; 8,465★ | Running Claude Code, Codex, and friends in local/private environments, especially where secrets must stay out of the sandbox and network policy must be dynamic |
| Kubernetes Agent Sandbox Kubernetes SIG Apps |
Sandbox, Claim, Template, and WarmPool CRDs plus an SDK; isolation comes from the RuntimeClass, usually gVisor/Kata | Kubernetes control-plane primitives | Official quickstart; the experimental path needs Docker/kubectl/KIND, production needs K8s plus RuntimeClass, networking, storage, and monitoring | Medium to high | Apache-2.0; 3,702★ | Mature K8s/SRE teams that want control-plane primitives rather than a full SaaS |
| Kimi AgentENV The Kimi (Moonshot AI) technical ecosystem |
A distributed Firecracker environment designed for agentic RL training; snapshots, forks, E2B-compatible API | Single machine or distributed training cluster | Official deployment docs; needs Linux kernel 6.8+ and KVM/PVM, multi-node also needs a gateway, scheduler, and shared storage | Medium to high Medium single-node; high multi-node |
MIT; 3,377★ | Model training teams doing large-scale RL rollouts that need to fork many sandboxes from identical state |
| BoxLite The BoxLite team and community |
A daemonless microVM runtime embeddable in an application, or runnable as a server; KVM on Linux, Hypervisor.framework on macOS | Single machine or embedded runtime | Getting Started; needs Apple Silicon macOS 12+, Linux with KVM enabled, or WSL2 with KVM support | Low | Apache-2.0; 2,291★ | Teams wanting strongly isolated VMs on one machine or inside an application process, without standing up Kubernetes first |
| Microsandbox Super Rad Company and the community |
A local-first microVM runtime, server, SDK, and MCP | Single machine or embedded runtime | Official quickstart; supports Apple Silicon macOS, KVM Linux, and Windows Hypervisor Platform; still beta | Low | Apache-2.0; 8,046★ | Single-machine/edge/local development that wants stronger isolation than Docker without standing up Kubernetes |
| AIO Sandbox The Agent Infra team and community |
One container combining browser, shell, files, MCP, VS Code Server, and VNC; by default just a Docker container | Docker capability image | Official deployment example; needs Docker; the official example uses seccomp=unconfined, so high-risk scenarios must layer on gVisor/Kata/VM |
Low Strong isolation must be added separately |
Apache-2.0; 5,822★ | Quickly building GUI/browser/computer-use environments and demos; do not treat the default Docker setup as a hostile multi-tenant boundary |
| Anthropic Sandbox Runtime Anthropic and the community |
Applies filesystem and network policy to any command; Seatbelt, bubblewrap, or a Windows restricted user + WFP | Local policy runtime | Installation and platform requirements; Linux needs bubblewrap/socat/ripgrep and usable user namespaces; Windows needs a one-time administrator install | Low | Apache-2.0; 5,102★ | Adding low-overhead protection to a CLI/agent on a developer machine; still shares the host kernel |
Whichever route you take, the gap between "I can start my first sandbox" and "this can carry untrusted production traffic" still has to be filled in: API authentication and TLS, deny-by-default egress, a secret broker, resource and concurrency quotas, image supply chain, audit logs, data/snapshot backup, version upgrades and rollback, plus real escape and exfiltration testing. A successful one-click install is not production readiness.
5. How Well-Known Agent Products Build This
This section applies the layered framework from part 1 to concrete products, focusing on harness, execution isolation, state, and public adoption.
5.1 Product Comparison
| Product | Harness / control plane | Actual execution and isolation | State and credentials | Public users/adoption | What a beginner should take away |
|---|---|---|---|---|---|
| OpenAI Codex (harness Apache-2.0, 120,596★) | An open-source harness manages context, tools, approvals, and sessions; the CLI, IDE, and app share the same core | Locally uses Seatbelt, bubblewrap/user namespaces, or the Windows native sandbox; Codex Cloud gives each task an isolated container | Secrets are removed during the agent phase; the network is off by default or goes through a proxy allowlist; results come back as a diff/PR | Cisco, GitHub, JetBrains | The local OS sandbox and the cloud container are two different modes; the open-source harness does not include the cloud control plane |
| Claude Code (main repository declares no standard OSS license, 143,631★) | A local/cloud harness maintains the loop and routes tool calls; Managed Agents can decouple from the customer's execution plane | The local Sandbox Runtime uses Seatbelt/bubblewrap plus a proxy to control files and domains; the web version gives each session a cloud sandbox | Locally, privilege escalation is approved by policy; the web version issues only scoped, short-lived capabilities and never hands down long-lived Git credentials | Accenture, Rakuten, Ramp | The main product cannot simply be labeled OSS; what is explicitly open source is the Sandbox Runtime |
| Devin (Cognition, closed-source product) | Cognition hosts the agent/harness; with Outposts, a customer-side orchestrator picks up tasks | Creates a Daytona sandbox per session, with the agent working over an outbound WebSocket | The sandbox follows session stop/resume/delete; the execution plane and the code can stay in the customer environment | The Devin Outposts architecture published by Daytona | The classic "closed-source SaaS harness + customer-controlled execution plane" |
| Pi (MIT, 100,264★) | A minimal, extensible agent loop; encourages replacing tools and UI via extensions | The local core has no permission prompts; read/write/edit/bash run with the current user's privileges by default. You can put all of Pi in a container, or use pi-sprites to route tool calls out to a Fly Sprite | Sprites can persist, checkpoint, and stop/resume; but with no network rules set, the whole internet is reachable by default, and tokens still have to be injected by the host per command or via a proxy | Fly's internal SpriteDoc and the Nous Research Hermes backend are public reference implementations | Pi is a harness, not a security boundary; its security depends on how the integrator replaces the execution tools or isolates the whole process |
| Manus (closed-source product) | The published 2025 architecture has a planner agent decomposing tasks and multiple executor agents collaborating, exposing about 27 tools to the model | Manus self-hosts E2B; each cloud computer is a Firecracker microVM containing Chromium, a terminal, and a filesystem | Sandboxes can live for hours and pause/resume; this is not one agent running a shell on a shared server | Manus is itself one of E2B's best-known public adoption cases | It is "multi-agent orchestration + a persistent cloud computer"; only the published 2025 architecture can be confirmed here, and the internal implementation may have evolved since |
| Lovable (closed-source product) | A hosted orchestration layer for AI app generation, building, and preview | Publicly used Modal Sandbox in 2025; Google later disclosed adoption of GKE Agent Sandbox / Agent Substrate | Build environments are isolated per task; credential and persistent-state details are not public | Public case studies from Modal and Google Cloud | The same product can use different execution substrates as it scales and its workload changes |
| Notion Workers (Notion, closed-source product) | Notion manages worker tasks, sessions, and tool orchestration | Each worker runs inside a Vercel Sandbox Firecracker microVM | A firewall and proxy constrain the network, injecting credentials at the egress as needed | Vercel's published customer case study | The agent never holds real secrets and can only reach external services through a controlled egress |
| DeepSeek Harness (MIT, 207,338★) | The Cordis kernel makes model/tool/skill/session/sandbox/storage/loop/scheduler/UI all pluggable; trajectories are an append-only event stream; ships Standard, Code, Minimal, and Creator modes | The local sandbox subsystem uses bubblewrap/Landlock, Seatbelt, or Windows ACLs/restricted tokens depending on the OS; it mainly constrains the filesystem and is still same-world/shared-kernel | Fails closed when a system capability is unavailable; a human can approve a wider mode. Containers, microVMs, and remote machines are meant to be implemented as swappable providers | Still a developer preview; as of this research, no verifiable well-known external production user was found | It is first of all a harness; a high star count does not mean its local process sandbox matches the multi-tenant isolation strength of E2B or CubeSandbox |
These products fall into three implementation routes:
- Local harness + OS sandbox: Codex CLI, Claude Code, DeepSeek Harness. Fast to start and natural to use, but still sharing the host kernel—good for "protecting the developer's machine," not for strong multi-tenant hosting of strangers' code.
- Hosted harness + a cloud environment per task: Codex Cloud, Claude Code Web, Manus, Notion Workers. The control plane owns the session, the execution plane is one container or microVM per task, and Git, network, and secrets are governed through a proxy.
- Harness decoupled from sandbox: Pi + Sprites, Open SWE + Daytona, Devin Outposts, Managed Agents. The tool protocol is the seam, so the execution location can be swapped without rewriting the agent loop—the most common direction for enterprise-owned compute and BYOC.
6. Choosing by Scenario
Building a First Demo
- First choice: E2B Hobby. Simple API, plenty of documentation and integrations, and self-hostable later.
- If you care more about fast creation, long-lived workspaces, or Windows/GPU: Daytona—accepting that the production platform is now closed source.
- For a coding agent on your own machine: Docker Sandboxes first; the OS sandboxes shipped with Codex/Claude/DeepSeek are a lighter option, as long as you understand they are not microVMs.
- If you only ever run your own trusted scripts, local Docker is a fine start; do not point it at untrusted public code.
Data Analysis / Code Interpreter
- General purpose: E2B.
- AWS enterprises: AgentCore Code Interpreter.
- GPU, Jupyter, batch Python: Modal.
- Agent already inside Cloud Run: Cloud Run Sandboxes.
- Azure ecosystem: Dynamic Sessions; watch the one-hour minimum billing granularity.
Coding Agents / Devin-Style Products
- One-off tasks at high concurrency: E2B, Daytona, Vercel.
- Keeping an environment around to resume after PR comments: Sprites, Runloop, Blaxel, Daytona.
- Bringing your own benchmark/eval: Runloop.
- Also hosting APIs, databases, and GPUs: Northflank.
RL, Evaluation, and Large-Scale Fan-Out
- Hosted: E2B, Runloop, Modal; Modal/Northflank/Daytona when GPUs are needed.
- Self-hosted general platform: OpenSandbox, CubeSandbox, or Kubernetes Agent Sandbox + gVisor/Kata/Firecracker.
- Self-hosted for training specifically: Kimi AgentENV; CoreWeave customers can look straight at CoreWeave Sandboxes.
- The capability that matters is not how fast a single sandbox starts, but template caching, snapshot forking, concurrency quotas, bulk creation rate, failure reclamation, and observability.
Browser / Computer Use
Browser sandboxes are an adjacent but distinct submarket: beyond code isolation, they have to manage Chrome, cookies/profiles, Playwright/CDP, VNC/screenshots, and anti-bot problems. Look at E2B Desktop, Daytona Computer Use, Runloop Browser/Computer, AWS AgentCore Browser, and AIO Sandbox. If your agent only drives web pages, a dedicated browser platform such as Browserbase is usually a better fit than a general-purpose Linux sandbox.
Strict Compliance, VPC, or On-Prem
- Fast hosted/BYOC: E2B Enterprise, Northflank BYOC, Daytona Customer-Managed Compute.
- Genuine self-control: OpenSandbox, CubeSandbox, OpenShell, or Kubernetes Agent Sandbox with gVisor/Kata/Firecracker underneath; BoxLite for single-machine embedding.
- AWS/GCP/Azure-native teams should consider AgentCore, GKE Agent Sandbox, and Azure Dynamic Sessions first—one fewer IAM, audit, and network stack to maintain.
7. The Easiest Security Mistakes to Make
A sandbox has to cover at least three lines of defense simultaneously:
- No escape: the agent must not read the host, kill host processes, or affect other tenants. This relies on microVM/gVisor/Kata, kernel patching, and layered isolation.
- No exfiltration or privilege abuse: even without an escape, reachable data must not be sent to arbitrary domains. This relies on deny-by-default networking, domain/CIDR allowlists, private networking, a secret broker, and short-lived credentials.
- No resource abuse: fork bombs, mining, unbounded downloads, filling the disk, and billing attacks. This relies on cgroup/VM quotas, process limits, disk quotas, timeouts, concurrency caps, and budget alerts.
A minimum pre-launch checklist:
- No internet by default; allow domains dynamically per task and forbid arbitrary outbound TCP/UDP.
- Never hand production API keys to the agent as environment variables; use proxy injection, short-lived tokens, or least-privilege accounts.
- One sandbox per user/task; never let multiple untrusted tenants share a writable filesystem.
- Set caps on CPU, RAM, disk, PIDs, file size, execution time, concurrency, and creation rate.
- Treat snapshots, templates, and uploaded files as untrusted input; enforce the same policy after a restore.
- Keep audit logs for commands, network, file changes, secret usage, and lifecycle events.
- Require authentication on port previews by default; share links should be short-lived, not public by default.
- Regularly test escapes, malicious package installs, DNS/HTTP exfiltration, fork bombs, filling the disk, and policy retention across pause/resume.
8. How to Run Your Own PoC Evaluation
Do not just compare vendor-claimed cold starts. Use your own workload, repeat each vendor at least 30 times, and record P50/P95:
- Empty sandbox creation through the first command completing;
- Creating from a custom template and installing dependencies;
- Cloning a mid-sized repository, running tests, starting a dev server;
- Pause/resume, snapshot/fork, failure retry;
- Creation success rate and rate limiting at 100 and 1,000 concurrent;
- The actual bill when the agent spends 90% of its time waiting on the model;
- Whether the default network, secret injection, and port exposure are safe;
- Idempotency when the API disconnects, the process keeps running, and the control plane retries;
- Whether data residency, logging, deletion, and snapshot retention meet compliance requirements;
- The same monthly workload with plan fees, storage, egress, concurrency, and ancillary services all included.
9. Final Recommendations
Absent existing constraints, narrow the field in this order:
- Fix the trust level first: trusted internal code can use containers; untrusted user/LLM code needs gVisor/Kata/microVM; strong multi-tenancy should prefer microVMs.
- Then the state model: ephemeral for second-scale tasks; pause/resume + snapshot/fork for coding and research agents; idle-to-zero + persistent disk for long-lived services.
- Then the deployment model: SaaS first for small teams; cloud-native for teams that already run a cloud platform; BYOC/self-hosting once regulation, data residency, or cost scale justify it.
- Compare prices last: use your real active-CPU ratio plus snapshot and network costs; never rank hourly rates that come from different billing models.
Notes on the Data
- All prices come from public vendor pages and may change by region, contract, resource type, and time; options without an independent public price are explicitly marked as preview or contact-sales rather than guessed at.
- "Startup speed" is as claimed by the vendor under inconsistent test conditions, and this article does not rank on it.
- The star snapshot is from 2026-09-01. Stars indicate attention only and do not equal security, activity, or production maturity. Daytona still has over 70,000 stars, but its production code has been closed source since June 2026; DeepSeek Harness's 200,000+ stars do not make its local process sandbox equivalent to a strongly isolated cloud platform.
- "Open source" is used strictly: a self-hostable full data plane and control plane is not the same thing as publishing only an SDK, CLI, or examples.
Appendix A: Core Terminology
| Term | Plain explanation |
|---|---|
| Harness / agent runtime | Owns the model loop, context, tool routing, approvals, and state; it decides "what to do" but provides no strong isolation by itself. |
| Sandbox | A restricted execution space or a whole execution service; you still have to confirm whether the substrate is a process, container, gVisor, Kata, or microVM. |
| OS sandbox | Restricts a local process using system mechanisms such as Seatbelt, bubblewrap, or restricted tokens; fast to start, but usually shares the host kernel. |
| Container | Uses namespaces, cgroups, and similar to provide an isolated view and resource limits; a mature ecosystem, but usually shares the host kernel. |
| gVisor | A user-space kernel project started by Google that reduces how much an application touches the host kernel directly; compatibility and I/O may cost something. |
| Kata Containers | A community project hosted by the OpenInfra Foundation that carries OCI/K8s containers inside lightweight VMs. |
| microVM / VM | Virtual machine isolation with its own guest kernel; microVMs are trimmed down for fast startup and high density, with Firecracker (AWS) as the canonical implementation. |
| KVM / nested virtualization | KVM is Linux's hardware virtualization capability; to run microVMs inside a cloud instance, confirm the instance exposes /dev/kvm. |
| Control plane / data plane | The control plane handles creation, policy, quotas, and lifecycle; the data plane actually runs commands and carries customer code. |
| Ephemeral / persistent / pause | Respectively: destroyed after the task, workspace retained, and compute suspended then resumed; confirm the state scope and billing during a pause separately. |
| Snapshot / fork | Saving environment state and cloning sandboxes from the same state; used for recovery, parallel attempts, and RL rollouts. |
| Template / OCI image | A base environment with the system and dependencies preinstalled; OCI support only means image compatibility, not that the substrate is an ordinary container. |
| Ingress / egress / allowlist | Inbound, outbound, and permitted lists; untrusted agents should have outbound traffic and preview ports restricted by default. |
| Secret broker / short-lived credentials | A trusted proxy injects or signs by policy so the agent never holds long-lived keys directly. |
| BYOC / self-hosted | BYOC brings execution resources into the customer's cloud; self-hosting maintains both the control plane and the execution plane. Neither means zero operational cost. |