← All posts
Guide

Harness engineering: the part of an agent you actually build

You cannot change the model much. The loop around it — which tools exist, what the model sees, what it may touch, where it runs — is yours, and the projects in this catalogue have already made those decisions in public.

An agent is a model inside an arrangement: a loop that calls it, a set of tools it may call back, a decision about what goes into its context, a rule about what it may touch without asking, and a boundary around where its code runs. That arrangement is the harness, and unlike the model it is entirely yours.

The distinction from a framework is worth keeping. A framework is a library you build with — LangGraph, Pydantic AI, the OpenAI Agents SDK — and choosing one is five questions about durability, language and maintenance. A harness is the running arrangement itself, and you have one whether or not you named it. Most teams discover theirs by accident, one tool and one bug at a time.

The vocabulary settled recently. Anthropic put “harness” in a title in November 2025, describing what lets an agent work across multiple context windows; LangChain reduced it to a formula in March 2026 — “Agent = Model + Harness”, where the harness is “every piece of code, configuration, and execution logic that isn’t the model itself”; and Birgitta Böckeler, writing on martinfowler.com in April 2026, splits it into guides that tell an agent how to proceed and sensors that tell it how the work is going.

The useful thing about 2026 is that these decisions are no longer folklore. Serious projects now ship them as configuration, document the reasoning, and — the part worth reading closely — state plainly where their own mechanisms stop working. Five decisions follow, each one readable off projects in this catalogue — and then the awkward question of whether any of it beats simply changing the model.

Decision 1: how many tools the model can see

The instinct is to connect everything. The projects that have lived with agents longest push the other way.

GitHub’s own MCP server groups its tools into twenty-three toolsets and turns on five of them by default: context, issues, pull requests, repositories and users. Its documentation gives the reason in one sentence — enabling only the toolsets you need “can help the LLM with tool choice and reduce the context size.” Both halves matter. Tool definitions are not free; they occupy the same context the work needs. And a model choosing among ninety-odd tools chooses worse than one choosing among a dozen.

The Chrome DevTools team draws the same conclusion harder. Their MCP server ships fifty-seven tools across eleven categories, keeps memory analysis and extension tooling switched off until asked, and offers a slim mode that collapses the entire surface to three: navigate, evaluate, screenshot. A vendor shipping a three-tool build of their own fifty-seven-tool product is telling you something about how tool count behaves in practice.

AWS’s suite arrives at the same place from the opposite direction: rather than one server exposing everything, it is sixty-odd small servers, and the intended use is to run the two or three that match today’s work.

The rule that falls out: treat the tool list as a budget, not an inventory. Decide per task, not per integration.

Decision 2: what the model sees, and when

Context is the other budget, and it is spent by the same things that make an agent useful: file contents, tool output, search results, the history of what already happened.

Three patterns in the catalogue are worth stealing. Context7 attacks stale knowledge rather than volume: it resolves a library name to an identifier, pins a version, and returns documentation for the release you actually run, so the model stops writing against the API it remembers from training. claude-mem attacks retrieval cost with staged disclosure — a search returns cheap index lines, a timeline places them, and only the records you open pay full price. Scrapling attacks input size at the door: its MCP server narrows a page by selector before the content ever reaches the model, and strips hidden DOM content that could carry instructions.

Lance Martin of LangChain names the three moves available here — reduce, by summarising older tool calls; offload, by keeping material outside the prompt until needed; isolate, by giving a token-heavy subtask its own agent and its own clean context. Each of the three projects above is one of those moves, sold as a product.

The shape they share is that context is filtered by the harness, deliberately, before the model is asked to be smart about it. The alternative — pour everything in and hope — is the most common way a working demo becomes an expensive, forgetful production system.

Decision 3: what it may do without asking

Here the honest projects are unusually blunt, and two quotes are worth putting side by side.

GitHub’s MCP server has a read-only mode that “acts as a strict security filter that takes precedence over any other configuration,” disabling write tools even when a configuration explicitly asks for them. It also has a lockdown mode for public repositories, which surfaces only content from users with push access — a defence against instructions smuggled into an issue by a stranger. And the documentation says this about it: “Lockdown mode remains a best-effort content filter, not a security boundary.”

DeerFlow says the same thing about its skill permissions, in its own words: “This is best-effort behavioral scoping, not a hard security boundary,” noting that instructions loaded through another tool are not captured.

Two large, careful projects, independently, telling you that the filter in front of the model is not the thing keeping you safe. That is the most useful sentence in agent engineering right now, and it points somewhere specific: the boundary has to live where the model cannot argue with it — in the credential and in the machine.

The credential half is concrete. AWS’s servers gate real changes behind explicit --allow-write and --allow-sensitive-data-access switches, per server, so an agent that reports on your account and an agent that can change it are different deployments. GitHub’s server has no repository setting at all: what the agent can reach is decided by the token you hand it. agentgateway is the version of this for organisations — a proxy that puts authentication, access rules and tracing in front of MCP and model traffic once, instead of each agent reimplementing them.

Decision 4: where the code actually runs

The moment an agent writes code and something runs it, the harness question becomes an infrastructure question.

E2B exists for this: isolated micro-VMs that start in well under a second, which is what makes a sandbox per task practical rather than aspirational. OpenHands makes the difference visible: run it from the published container image and the agent reaches the project directory you mounted, while the npm and from-source launchers carry the project’s own warning that the agent server has full access to your filesystem. DeerFlow makes the choice explicit at deploy time — local process, Docker, Kubernetes or E2B — which turns “should the agent run code” into a configuration line rather than a matter of trust.

The counter-example is instructive rather than damning. OpenClaw is built for a single operator on their own machine, and its README states that tools run on the host unless you configure sandboxing. That is a coherent design for one person’s laptop and a bad one for a shared deployment — which the project also says. Read the sandbox default before the feature list; it tells you which of those two the project is.

Decision 5: whether the loop checks its own work

A harness that produces work nobody verified produces plausible work. Plausible is the failure mode specific to this technology, and the cheapest defence is a step in the loop that can say no.

Aider wires the obvious one: it runs your linters and test suite after each change it makes and repairs what they report. The check is not clever, and that is the point — it is a fact about the code rather than an opinion about it. Cline puts a different check in a different place: a plan mode that explores and proposes before an act mode executes, with each edit and command gated by approval and recorded as a reviewable diff you can roll back.

Those are two answers to the same question — is the verifier a machine or a person — and the honest guidance is to have the machine one before you need the human one. Tests, types and linters are the parts of your codebase that already disagree with wrong work; an agent gets a compounding benefit from them that a human reviewer never quite gets, because it can act on the disagreement immediately and try again.

Does any of this beat changing the model?

The claim you will hear is that the harness matters more than the model. The evidence is more interesting than the slogan, and it points at a different reason to do the work.

The strongest published gain comes from LangChain, in February 2026: holding the model fixed at GPT-5.2-Codex, harness changes moved their agent from 52.8 to 66.5 on Terminal-Bench 2.0. It is a vendor measuring its own product, and the caveat is theirs too — a run with Claude Opus 4.6 scored 59.6 on an earlier harness version, which they attribute to not having repeated the same tuning loop for that model. Harness gains, in their own telling, are tuned per model rather than won once.

A position paper from May 2026 argues the general case, that “harness-induced variance can substantially exceed model-induced variance, including cases of model ranking reversal,” and asks that benchmark results disclose the harness at all. Worth reading as an argument about how we measure, not as a measurement.

The most careful measurement I found says something quieter and more useful. Vats and Golev ran two models across three open-source harnesses — Goose, OpenCode and the OpenHands SDK — on a fifty-task benchmark subset in June 2026. Accuracy barely moved: pass-rate differences of zero to eight points, with confidence intervals including zero except at the widest gap. Cost moved enormously: up to a fortyfold difference in tokens per solved task. Failure modes were harness-specific and repeated across both models.

Take that seriously and the case for harness engineering changes shape. Its measured effect is not that your agent becomes cleverer. It is that the same work costs forty times less or forty times more, fails in a way you recognise or a way you cannot see, and touches a blast radius you chose or one you inherited.

That also explains what to build and what to skip. The sceptics have a point worth keeping: Han Lee argues that most harness code “is going to dissolve into the next generation of models,” and Lance Martin, who works on this for a living, says the job over time is to strip structure away as models improve. They are describing the scaffolding that compensates for model weakness — the elaborate prompt chains, the hand-held decomposition. Delete-on-upgrade is the correct expectation for that half. The other half — which credential the agent holds, where its code runs, what enters the context, whether tests ran — is not compensating for anything. No model upgrade retires the question of what the agent is allowed to touch.

What this means for your next month

Order the work by which mistake is most expensive to make.

Start with the boundary, not the capability. Decide what credential the agent holds and where its code runs before you decide what it can do. Those two facts bound every future mistake, and both quotes in this article exist because two large projects wanted to be clear that the filter in front of the model does not.

Then cut the tool list. If you connected an MCP server and left every toolset on, you are paying context for tools the model will not use and making its choices harder. GitHub’s five-of-twenty-three default is a reasonable model for your own surface.

Then decide what fills the context, deliberately. Something has to be filtered by the harness, and the choice is whether you designed the filter or inherited it from whatever the last tool happened to return.

Then add the verifier. If the work has tests, run them in the loop. If it does not, the first agent-shaped project to take on is the one that gives it tests.

Do not start by changing frameworks. That is the move that feels like progress and changes the least: the framework decides how you express the loop, while everything above decides whether the loop is safe, affordable and correct. If you are choosing one anyway, five questions narrows the field faster than a feature table.

Projects in this post