Coding Agents
Agents that read, write and run code against a real repository. Benchmark scores are a starting point, not an answer: sandboxing, token cost on large codebases and how reviewable the output is matter more day to day.
22 projects
The distinguishing idea is that everything is declared rather than hard-wired: you define named agents in JSON or a markdown file, give each one its own model and its own per-tool permissions — allow, ask or deny for reading, editing and running commands — and switch between them with a keystroke. Two built-in primary agents ship this way, one with full access and one that asks before editing or running anything. It is MIT and provider-agnostic, but it moves fast and it recently changed hands, so old links and pinned versions go stale quickly.
It works the way the other terminal agents do — read the repository, propose changes, run commands — but the boundary it enforces is unusual: on macOS and Linux it uses operating-system sandboxing to restrict which files and commands the agent can reach, so the limit does not depend on the model choosing to respect it. Sign-in runs through a ChatGPT plan or an OpenAI API key, which is the trade: the code is Apache-2.0, the model behind it is not the open part.
Ask an agent for a date picker and you can get a dependency, a wrapper component, a stylesheet and a conversation about time zones, when the browser has had `<input type="date">` for years. Ponytail is a single skill aimed at that habit: it pushes the agent towards the smallest thing that works, and towards not writing anything when the platform already provides it. What makes it worth a look rather than a joke is that the maintainers measured it — twelve real tasks on a real repository, with and without the skill — and published the method alongside the numbers, including a note that an earlier, more flattering figure was the per-task ceiling rather than the average.
Google's own agent for the terminal: it reads a repository, edits files, runs shell commands, and can ground an answer in a live Google Search without extra setup. A read-only plan mode drafts the work before anything is written, a policy file settles in advance which tools may run, and one extension can install MCP servers, commands, subagents, skills and hooks together. It talks only to Gemini models, so it is a poor fit if you expect to switch providers or work offline — and since 18 June 2026 a personal Google sign-in no longer reaches it, leaving an API key, Vertex AI, or a Code Assist Standard or Enterprise licence.
Left alone, a coding agent starts writing code the moment you finish the sentence. This is a set of skills that puts the usual engineering order back: nine commands cover the stages from working out what to build to shipping it, and the skills behind them attach on their own when the work calls for one — designing an interface pulls in the interface skill, touching the front end pulls in that one. Notably the automatic mode does not remove the checks, only the person stepping between tasks: each task is still test-driven, committed on its own, and paused on failure. It is Markdown, so it installs into around seventy agents through the open skills command.
Edits files, runs tests and browses documentation inside an isolated container, so a failed run cannot damage the host. Competitive on SWE-bench style benchmarks, but budget carefully before pointing it at a real repository — long autonomous runs consume a great deal of tokens and still need review before anything is merged.
Hand MetaGPT a one-line requirement and a team of LLM roles works through it in turn, leaving a requirements document, a system design and source files in a folder under `workspace/`. Which roles show up depends on where you installed from, and the gap is not cosmetic: the published 0.8.2 package hires a product manager, an architect, a project manager and an engineer, while current main — the README's clone-and-install route — hires a team leader, a product manager, an architect, an engineer and a data analyst, with the project manager commented out. Under the demo sits a general framework: you subclass Role and Action, declare which upstream actions each role watches, and a shared Environment routes the messages that put the team in order. A separate agent, the Data Interpreter, does data work alone, planning the steps, writing Python and retrying what fails. The catch is scope — MetaGPT's own FAQ says functions start being left unimplemented past roughly 500 lines of generated code.
Type `i` in a project directory and it reads your files, edits them and runs commands inside a sandbox you configure. What sets it apart is `/harness`: it swaps the whole model-facing surface — system prompt, tool schema, message conversion — for the one a given vendor tuned its model on, which is how it aims to get usable work out of inexpensive models like DeepSeek or Kimi. The trade-off is what you are actually adopting: this is a fork of OpenAI's Codex CLI, rewritten in Rust and still on the 0.0.x releases it started in June 2026, sharing little more than a name and a repository with the Python Open Interpreter most people have heard of.
Runs in VS Code and asks for approval before each file write or terminal command, which keeps a human in the loop without breaking flow. MCP support lets it reach tools beyond the editor. The approval prompts are the safety mechanism — turning them all off removes the main reason to prefer it.
Runs locally as a CLI or desktop app and gains capabilities through MCP extensions rather than a fixed tool list, so the same agent handles code, infrastructure and third-party services. Model-agnostic, which matters if you expect to switch providers.
Every change lands as a git commit, which makes an AI edit exactly as reviewable and revertible as a colleague's. Its repository map lets it work on codebases far larger than a context window. Terminal-only by design — a strength if you live there, a non-starter if you want it inside an editor.
Autocomplete, chat and edits driven by whichever model you point it at, including one running on your own hardware. That makes it one of the few realistic options where code cannot leave the building. Configuration is more involved than the hosted alternatives — the price of that flexibility.
The central idea is the CodeAgent: rather than returning a structured tool call, the model writes a short Python snippet that calls your tools directly, so nesting a call inside a loop or a conditional comes from the language itself instead of from framework machinery. The whole agent loop is about a thousand lines, which is both the appeal and the ceiling — durable state, workflow definitions and persistence are things you build yourself. A conventional ToolCallingAgent is included for the JSON style where that fits better.
It installs as one Go binary from almost any package manager and runs on an unusually wide set of systems, including Windows PowerShell, WSL, Android and the BSDs. Inside, sessions are per project, the model can be swapped mid-conversation without losing context, and language servers feed it the same code intelligence an editor would. The licence is the thing to read first: it is not OSI-approved, it forbids offering a competing product, and each release only becomes MIT two years after it ships.
A long task fails in a particular way: the agent had a plan, the context was cleared or compacted, and what comes back is an agent confidently finishing a different job. This skill moves the plan out of the conversation and onto disk — one file for the plan, one for what has been found, one for what is done — and puts them back in front of the agent on every turn. Because the plan is a file, it outlives a clear, a crash and a compaction, and you can read it yourself while the run is going. An optional gate stops the agent declaring completion until the plan says so. The maintainers publish blind A/B results, which is more than most skills offer.
Devika turns a one-line objective into files on disk through a fixed chain of specialists: a planner breaks it into a numbered checklist of steps, a researcher condenses that into at most three search queries, a browser opens the first hit for each one and keeps the text, and a coder writes the result into a project folder. The interface is the reason to look: the plan, the agent's running commentary, a screenshot of the page it is reading, the shell output and the token count are all on one screen, so a run that goes wrong usually shows you the step where it did. What it never did was finish. The README's stated ambition was to match Devin's score on the SWE-bench benchmark, and the two files in the repository meant to hold that result are a placeholder line and an empty file. With the code untouched since 2024, Devika is worth reading as the clearest surviving example of that year's open-source Devin wave, not as something to build on.
`npx @openai/codex-security scan .` walks a checkout and leaves JSON results on stdout, while `--mode deep` spreads discovery across workers and subagents until it stops turning up anything new. Across runs, `scans compare BEFORE_SCAN_ID AFTER_SCAN_ID` matches findings by root cause and labels them new, persisting, reopened, resolved or unknown, so a second scan reports movement instead of restating the whole report. It is a poor fit as a pre-commit gate: deep discovery runs until `--max-time-hours`, which defaults to 96. Access is also gated — the CLI needs access to Codex Security, and some cybersecurity requests and protected findings require Trusted Access for Cyber approval, so cloning the repo on its own scans nothing.
Cloudflare OS pairs an agent chat with "Gadgets" — small apps the built-in coding agent writes for you, each running as a Dynamic Worker with its internet access disabled, its client confined to a sandboxed iframe that reaches the server only over Cap'n Web via `postMessage()`. External services are reached through Gatekeepers, per-service Workers that wrap an API in Cap'n Web, log every call, and simulate side-effecting actions so the agent keeps queueing work while approvals wait for you to review them in bulk. The commitment to Workers runs deep: every workspace is a Durable Object, every Gadget a Dynamic Worker Facet, and Facets and Dynamic Workers were added to the runtime specifically for this project — so hosting means a Cloudflare account, and the deploy-to-your-own-servers-on-`workerd` path is still marked COMING SOON in the README. The maintainers call the August 2026 v2 rewrite an early access release with many rough edges, and many Gatekeepers need configuration of their own — including OAuth client credentials per service — before GitHub or Google will connect.
A coding agent working on an iOS project can write Swift all day and never find out whether it compiles, because the commands that would tell it are long, particular and easy to get subtly wrong. This exposes them properly: build a scheme, boot a simulator, install and launch the app, read the logs, run the tests. Because the agent gets a result it can act on, the loop closes — it fixes the error it just caused instead of handing you a diff to try. It ships as both an MCP server and a plain command-line tool, with optional skills that tell the agent how to use whichever one you installed. It is macOS-only and tracks a current Xcode.
The expensive habit of a coding agent is grep, then read the eight files it matched, then read three more because the answer was not there. Semble replaces that with one question in plain language and a handful of relevant snippets back — the maintainers put the saving at roughly 99% of the tokens against grep plus reading. It runs entirely on the machine, on CPU, with no API key and no service to call, and indexing a full repository takes under a second, which is what makes it practical to point at whatever the agent happens to be working on. Installation detects the agents you have and offers three ways in: an MCP tool, a line in your agent instructions, or a dedicated search sub-agent.
A task gives the model a codebase and an issue and asks for a patch; the patch is accepted only if the tests that were failing now pass and the ones that were passing still do. Using the projects' own test suites as the judge is what makes the result mean something concrete rather than resemble a rubric score. What it measures is correspondingly narrow: producing a fix for a problem someone else already found, filed and described — not deciding what to build, not reviewing, not navigating a codebase without an issue to anchor on.
Rather than a fixed set that saturates and stops distinguishing anything, this is a continuous benchmark: tasks are added and fixed over time and the dataset is published as tagged releases, run by a separate harness called Harbor. That design is the strength and the complication — a number without a dataset version attached does not mean much, version 1.0 lives in its own repository that the original URL now redirects to, and the documented first step is running the reference solutions repeatedly to confirm your own sandbox is sound before trusting any result.