Harness engineering: everything around the model
What harness engineering is, who named it, how to do it, and why an agent with a good harness beats a better model with a bad one.

Where it came from
Anthropic was already using the word harness in November 2025. In “Effective harnesses for long-running agents” they called the Claude Agent SDK “a general-purpose agent harness.” They also showed how they got Opus 4.5 to make progress across many context windows. An initializer agent sets up the environment on the first run. A coding agent then moves forward one feature per session and leaves artifacts for the next session. The problem they were solving: every session starts with no memory, like a shift of engineers who don’t know what the previous shift did.
Addy Osmani credits Viv Trivedy with naming harness engineering as a discipline. Trivedy posted “Anatomy of an Agent Harness” on X with the formula “Agent = Model + Harness. If you’re not the model, you’re the harness.” Around the same time Birgitta Böckeler wrote about it on martinfowler.com, and some people credit her with the term itself. The Habitat-Thinking doc, for example, ties the term directly to the classic test harness: tests don’t make code correct, they detect when it stopped being correct.
OpenAI is the one that made it a word everyone uses, with the post “Harness engineering: Leveraging Codex in an agent-first world.” They built an internal product with no hand-written code at all. Over five months, about a million lines were written and roughly 1,500 PRs were merged. The team started with three engineers and averaged 3.5 PRs per engineer per day. Their key line: “Humans steer. Agents execute.” From there it only spread. In March 2026 Anthropic published a planner, generator and evaluator architecture. In July Lilian Weng wrote about the harness as a central layer in recursive self-improvement. And this month Marmelab surveyed the field through 246 repos and 57 publications.
What it actually is
Harness engineering is the design of everything that wraps the model inside an agent. The prompt and AGENTS.md, the tools, the context policy, hooks, sandbox and permissions, subagents, verification loops, and recovery paths for when something breaks. The cleanest definition is Lilian Weng’s: the system around the base model that determines how it plans, calls tools, sees and manages context, stores artifacts and evaluates results. The practical version is Osmani’s: every time the agent makes a mistake, you build something that stops it from making that mistake again.
What it isn’t. It isn’t prompt engineering, because the prompt is one component out of ten. It isn’t context engineering either, which is about what the model sees. The harness includes that, and adds what the model can touch, how it gets checked and when it gets reset. And it isn’t a framework. An Agent SDK or LangChain are building materials. The harness is what you assemble from them for your repo and your team.

How to do it
- A map, not an encyclopedia. A short AGENTS.md or CLAUDE.md that points to docs. The agent simply won’t read a 2,000-line doc to the end.
- Decomposition and artifacts between sessions. A feature list in JSON, a progress file, and a commit after every step. This is what stopped Anthropic’s agent from trying to build the whole app at once and getting stuck halfway with no context.
- Deterministic checks after every action. Tests, lint, schema and type-check running through hooks. Not “please run the tests” in the prompt.
- An evaluator separate from the agent that writes. Anthropic split the generator from the evaluator, inspired by GANs, because an agent grading its own work tends to declare victory too early.
- Every mistake becomes a rule. The agent broke a convention? Now there’s a lint rule or a CI check, not a comment in the chat.
- Permissions and sandbox. Decide up front what the agent is allowed to delete, push and deploy.
- Observability. Logs that both the agent and a human can read. Without them you’re debugging guesses.
The shortest example of item 3, in Claude Code: a hook that runs the tests after every edit, where exit 2 sends the failure back to the agent.
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [{ "type": "command", "command": "npm test --silent || exit 2" }]
}
]
}
}

What it buys you, and what it doesn’t
There is one number that was actually measured. Someone ran the same model through 8 different harnesses on the same 25 tasks, with the same provider and the same tools. Success rates ranged from 68% to 88%. Marmelab, who cited the experiment, admit it’s about the only thing anyone has measured properly. Everything else, including most of what’s written here, is field experience, not a controlled experiment.
What it doesn’t buy you. A harness doesn’t make the model smarter, as the walkinglabs course puts it. It gives the model a working system with a closed loop. And it costs money and time. Every rule you add is code you have to maintain. Rules pile up and conflict, and eventually you end up with a harness nobody understands.
Then there’s build vs buy. Rajit Khanna of prismvideos deleted an agent he had built on the Vercel AI SDK and moved to Hermes, because memory, skills and automations came ready-made there. His advice: if someone has already built the generic part, don’t build your own harness. Invest in the tools and knowledge only you have.

From our own pipeline
Every morning aibriefing.dev runs a graph of about ten agents with a merger step. After publish, a QA evaluator runs and flags P0s. Our wrapper around claude -p is a small harness in every sense. It retries transient failures and repairs broken JSON before anyone tries to parse it. That closed three production incidents. A better model didn’t help. A loop of a few dozen lines did. The main lesson: our worst bugs weren’t wrong answers. They were silent failures, a run that finished “successfully” with empty output. That’s why the wrapper lives in shared/ and isn’t copied into each agent. When this logic was spread across several copies, fixing one of them left the others broken.
Sources
- Harness design for long-running application development
- The State Of AI Harness Engineering 2026
- Effective harnesses for long-running agents
- Harness engineering: Leveraging Codex in an agent-first world · 2026-06-05
- Harness engineering for self-improvement · 2026-08-04
- Learn Harness Engineering · 2026-05-18
- Harness Engineering · 2026-08-27
- Building agents without harness engineering · 2026-06-11
- Agent Harness Engineering · 2026-07-12
