Sandboxes & Runtimes

Isolated environments where agent-generated code can run without endangering the host. A prerequisite, not an optional extra, once an agent executes anything it wrote itself.

6 projects

deepseek-harnessFrameworks

`dsh` is the agent harness DeepSeek AI develops itself. `npx @deepseek-ai/dsh web` brings up a Web UI on `http://127.0.0.1:3080`, and the architecture behind it is one where everything is a plugin, powered by the Cordis runtime, so it suits people who intend to extend a harness rather than only operate one. The caveat comes from the project itself: it is labelled a developer preview and warns in capitals that there will be compatibility-breaking changes, so plugins written today should budget for rework. It is also a poor fit if you want a documented scripted or embedded entry point right now, since the README covers only the `web` command and a source checkout, and links out to a development guide and architecture documentation without saying what either contains.

OfficialMIT202k+21.4k
OpenHandsCoding Agents

Edits files, runs tests and browses documentation inside an isolated container, so a failed run cannot damage the host. Competitive on SWE-bench style benchmarks, but budget carefully before pointing it at a real repository — long autonomous runs consume a great deal of tokens and still need review before anything is merged.

OfficialMIT85.5k+746
CuaComputer Use

Computer use runs into the same wall every time — the agent needs a machine, and it cannot be the one you are working on. Cua is four answers to that, which is worth knowing before you install it, because you almost certainly want one of them rather than all four. The drivers let an agent click and type in native applications in the background, without taking the cursor or the focus away from you. The sandbox gives Linux, macOS, Windows or Android behind a single API, locally under QEMU or in the maintainers' cloud. Cua-Bench runs OSWorld and similar evaluations and exports trajectories for training. Lume creates macOS virtual machines on Apple Silicon using Apple's own virtualisation.

MIT21.9k
E2BSandboxes

Starts an isolated micro-VM in well under a second, which is what makes per-task sandboxing practical rather than theoretical. If an agent executes code it wrote itself, something like this is a requirement, not a nicety. Self-hosting is possible but meaningfully more work than the hosted path.

OfficialApache-2.013.6k+68
cloudflare-osSandboxes

Cloudflare OS pairs an agent chat with "Gadgets" — small apps the built-in coding agent writes for you, each running as a Dynamic Worker with its internet access disabled, its client confined to a sandboxed iframe that reaches the server only over Cap'n Web via `postMessage()`. External services are reached through Gatekeepers, per-service Workers that wrap an API in Cap'n Web, log every call, and simulate side-effecting actions so the agent keeps queueing work while approvals wait for you to review them in bulk. The commitment to Workers runs deep: every workspace is a Durable Object, every Gadget a Dynamic Worker Facet, and Facets and Dynamic Workers were added to the runtime specifically for this project — so hosting means a Cloudflare account, and the deploy-to-your-own-servers-on-`workerd` path is still marked COMING SOON in the README. The maintainers call the August 2026 v2 rewrite an early access release with many rough edges, and many Gatekeepers need configuration of their own — including OAuth client credentials per service — before GitHub or Google will connect.

OfficialApache-2.09.3k+623
Terminal-BenchBenchmarks

Rather than a fixed set that saturates and stops distinguishing anything, this is a continuous benchmark: tasks are added and fixed over time and the dataset is published as tagged releases, run by a separate harness called Harbor. That design is the strength and the complication — a number without a dataset version attached does not mean much, version 1.0 lives in its own repository that the original URL now redirects to, and the documented first step is running the reference solutions repeatedly to confirm your own sandbox is sound before trusting any result.

OfficialApache-2.0556+32