- An agent is a model plus a harness: the loop, tools, context assembly, state, permissions and checks that surround the weights. Teams showed in 2026 that changing only the harness can move an agent from mid-table to the top of a benchmark, so the harness is now where reliability, and much of the cost, is won or lost.
- Harnesses work through two kinds of control: guides that steer the model before it acts (system prompts, AGENTS.md files, skills, specs, type definitions) and sensors that catch failure after it acts (compilers, tests, linters, validators, evaluators). The discipline is turning every observed failure into a new guide or sensor so it cannot recur silently.
- Treat the harness as a product: version it, evaluate model and harness as a pair, measure cost per completed task, and keep its guides and sensors portable across models. Models are rented and swapped often; the harness is the part of the stack you actually own.
Agent equals model plus harness
The way enterprises talk about getting good behavior from language models has gone through three phases in roughly three years. First came prompt engineering, the belief that quality was mostly a wording problem. Then came context engineering, the recognition that what enters the window matters more than how you phrase the request, covered in the context engineering playbook. In early 2026 a third phase got its name: harness engineering, the design of the entire machine that surrounds the model in an agent system.
Two publications crystallized the idea. Birgitta Böckeler's essay on martinfowler.com gave practitioners a working vocabulary, sorting harness components into guides that steer an agent and sensors that check it. OpenAI described how one of its teams built an internal product of roughly a million lines of code in which, by design, no line was typed by a human: around 1,500 pull requests merged, with three engineers steering coding agents rather than writing code. The interesting part of that account was not the model. It was the scaffolding the team kept adding whenever an agent struggled: better documentation for the agent to read, stricter linters, new tests, clearer architectural rules.
The evidence that the harness, and not only the model, sets the ceiling is now hard to ignore. LangChain reported moving its coding agent from around thirtieth to fifth place on the Terminal Bench 2.0 leaderboard by changing only its harness. Several vendors have reported double-digit reductions in token spend from rebuilt harnesses at flat or better accuracy. Within a capability tier, models are increasingly interchangeable; whether that capability turns into completed, correct work is decided by everything around them.
A working definition: the harness is everything in an agent system that is not the model weights. That includes the control loop, the tool layer and its permissions, context assembly and compaction, state and memory, verification, recovery and telemetry. It is distinct from a framework, which is a library you might build a harness with, and from a platform, which is where harnesses run. Teams that say "we use framework X" have usually chosen a starting harness, not finished one.
Anatomy of a harness
Harnesses differ wildly in sophistication, but the components recur. Each exists because a specific, recognizable failure happens without it:
| Component | What it does | Typical failure when missing |
|---|---|---|
| Control loop | Decides when to call the model, run a tool, stop or ask a human | Runaway loops; premature claims of being done |
| Guides | Instructions, conventions and examples loaded before each step | Inconsistent output; the agent reinvents conventions every run |
| Tools and permissions | Typed actions with scoped credentials | Over-broad access; the confused-deputy problem |
| Context assembly | Selects, orders and compacts what enters the window | Context rot; instructions lost deep in long runs |
| State | Plans, scratchpads, task lists and checkpoints kept outside the window | Repeated work; no way to resume after a crash |
| Sensors | Deterministic checks and evaluators run on each output | Silent errors reach a human, or production |
| Recovery | Retries, rollback and escalation paths | One failed step ends a long task |
| Telemetry | Traces of every step, tool call, token and cost | Nothing to debug and nothing to improve |
Not every system needs all eight. A single-turn extraction call needs little more than a schema and a validator. A coding agent that works unattended for an hour needs every row, and the rows interact: state makes recovery possible, telemetry makes sensors improvable, and permissions bound what a misbehaving loop can do. The shape that emerges is consistent enough to draw:
GUIDES (feedforward)
system prompt | AGENTS.md | skills | specs | types
|
v
+---------------------------------------------+
| CONTROL LOOP |
| assemble context --> MODEL --> action |
| ^ | |
| | v |
| state / memory <------- tools (scoped) |
+---------------------------------------------+
|
v
SENSORS (feedback)
compiler | tests | linters | schema | evaluator
|
pass --------+-------- fail
| |
deliver retry, with the
failure as context
The figure makes the central design question visible. Every failure the system will ever have either gets prevented upstream by a guide, caught downstream by a sensor, or escapes to a person. Harness engineering is the work of moving failures from the third category into the first two.
Guides: steering before the step
Guides are feedforward controls: they raise the probability that the model gets it right on the first attempt. The familiar ones are the system prompt and few-shot examples. The ones that matter most in 2026 live outside the prompt. Repository instruction files such as AGENTS.md, now an open convention stewarded by the Agentic AI Foundation alongside MCP, tell coding agents how a codebase is built, tested and organized. Agent skills package procedures an agent loads on demand. Specifications, the subject of the spec-driven development deep-dive, state intent precisely enough to check. Type definitions, schemas and architecture decision records steer quietly but powerfully, because they constrain what can be produced at all.
Four properties separate guides that work from guides that decorate:
- Layered and loaded on demand. Dumping every convention into every request spends the context budget and buries the instruction that matters. Load a short core always and the rest when the task calls for it.
- Specific and observable. "Run the test suite and paste the result before claiming completion" beats "be careful". An instruction stated as an observable condition can be paired with a sensor; a virtue cannot.
- Close to the work. Guides that live in the repository next to the code they govern, versioned with it, stay accurate. Guides in a wiki drift until the agent faithfully follows rules the team abandoned a year ago.
- Owned. Every guide has a maintainer. A stale guide is worse than none, because an agent executes it with complete diligence.
Sensors: catching failure after the step
Sensors are feedback controls: they observe what the agent produced and let it correct itself before anyone reviews the work. They form a ladder, and the order of the rungs matters. Deterministic checks come first: compilers, type checkers, unit and integration tests, linters, schema validators, policy checks, limits on how large a change may be. They are cheap, fast and unambiguous, and an agent can read their output and act on it. Model-based checks come second: an evaluator model scoring output against a rubric, or a critique pass looking for what deterministic checks cannot see, using the methods in the evaluation methods deep-dive. Humans come last and are reserved for what matters most, as the human-in-the-loop deep-dive describes.
The mechanism that makes sensors powerful is the loop back. When a test fails, the harness hands the failure message to the model as new context and lets it repair the work, with a cap on retries so a hopeless task fails fast instead of burning budget. A sensor that only reports to a dashboard is monitoring. A sensor that feeds the loop is engineering.
Over time the most valuable habit is the ratchet: every failure observed in review or production becomes either a new guide, so it is prevented, or a new sensor, so it is detected. The OpenAI team's account reads like a long sequence of these moves. When an agent got something wrong, the question was not how to reword the prompt but which capability, document or check was missing. Run that ratchet for a quarter and whole classes of error become structurally impossible to repeat.
Sensor quality deserves the same seriousness as code quality. A flaky test teaches the loop that failures are noise and can be retried away. An evaluator with a high false-positive rate trains the team to ignore it. Measure each sensor's catch rate and false-alarm rate, and retire the ones that do not earn their latency.
The harness is the unit you evaluate
If the harness moves results as much as the model, then a model leaderboard is measuring something slightly different from what you will deploy: someone else's model paired with someone else's harness. The useful unit of evaluation is the pair, your model inside your harness on your tasks, which is exactly the posture of the evaluating agents deep-dive. Trajectory-level evaluation matters more here than answer-level scoring, because harness defects show up in the path: wasted tool calls, loops, abandoned plans.
Harness changes are code changes and need regression protection like any other. A rewording of a guide, a new compaction rule or a changed retry threshold can lift one task family and quietly break another, so harness releases belong in the same evaluation-gated pipeline described in the versioning and regression deep-dive.
Model swaps deserve explicit budgeting. Harnesses accumulate tuning to one model's habits: how it reads tool descriptions, how it responds to failure messages, how much guidance it needs. A newer, stronger model can perform worse in an old harness until the guides and thresholds are re-tuned. Keep model-specific adjustments in one place, route calls through a model gateway, and re-run the full evaluation suite on every swap.
The metrics that matter follow from the definition. Track task completion rate rather than answer quality alone; cost per completed task rather than cost per token, the framing of the LLM economics deep-dive; human interventions per task; the share of failures caught by sensors before a person saw them; and time to completion. A harness improvement that raises completion while cutting interventions is worth more than a model upgrade that does neither.
Build, borrow or buy
Nobody should build a harness from nothing in 2026. Model vendors' agent SDKs and the open frameworks surveyed in the framework landscape deep-dive ship reference harnesses with a working loop, tool plumbing, compaction and permission prompts. Coding agents are harnesses too, configurable through instruction files, hooks and permission settings, as the coding agents deep-dive describes. Borrow the generic machinery.
Build the parts that encode your enterprise. Your guides carry your conventions, policies and domain procedures. Your sensors carry your test suites, your compliance checks and your definition of correct. Your tools and their permissions carry your systems and your risk appetite. These are the differentiated, durable assets, and they should be owned by named teams and stored as portable artifacts: files, scripts and evaluation suites that can move from one harness to another.
At scale, offer a paved road. A platform team can provide the shared sensors every agent needs, such as secret scanning, PII detection and policy checks; a standard permission model; shared telemetry; and templates for guides. Product teams then spend their effort on domain guides and domain sensors, which is where their knowledge actually is. Avoid deep forks of a vendor harness: the vendor will keep improving the generic machinery, and a fork turns every upgrade into a merge project.
The architect view
The strategic point is simple. Models are rented, frequently replaced and broadly comparable within a tier. The harness is where an enterprise's knowledge, controls and reliability actually live, and it is the part of the agent stack that compounds. An organization that invests in its harnesses gets better with every model generation. One that invests only in model selection starts over each time.
Four commitments turn that into practice. First, inventory the guides and sensors every production agent depends on, and version them in repositories next to the systems they govern. Second, build the sensor ladder deliberately: deterministic checks before model-based ones, model-based before human review. Third, institute the ratchet: every incident or rejected output yields a new guide or sensor, tracked to closure like a postmortem action. Fourth, evaluate model and harness as a pair, and keep the harness portable so the next model swap is a re-tune rather than a rebuild.
The next model will arrive before your next planning cycle. The harness you build this quarter is what will still be standing when it does, and it is what will decide whether that model's extra capability shows up as finished work or as a more articulate way of being wrong.