
Harness Engineering: The Missing Engineering Layer Between LLMs and Production
LLMs provide probabilistic intelligence, but production software needs predictable behavior. Learn how harness engineering makes AI agents reliable with state, tool constraints, validation, evals, and observability.
Harness Engineering: The Missing Engineering Layer Between LLMs and Production
An agent can complete a task successfully ten times in a demo. Production engineering asks a different question:
What happens on the 11th run when the tool fails, the model produces an invalid argument, the context becomes too large, the API times out, the agent enters a loop, or the cost exceeds the allowed budget?
That is where most “working” AI prototypes stop being systems and start becoming incidents.
The central idea of this article is simple:
The model is not the system.
A large language model can reason, call tools, write code, and choose among possible actions. It provides probabilistic intelligence. It does not, by itself, provide durable state, bounded execution, authorization, recovery, auditability, or a definition of “done.” Those properties have to be engineered around it.
That surrounding layer is what I mean by harness engineering: the design of the runtime, constraints, feedback loops, and operational controls that turn a capable model into a dependable production component.
Why a successful demo fails in production
A demo usually tests the happy path. The prompt is carefully chosen, the tools are available, the context is small, the network is healthy, and a human is watching closely enough to restart the run when something looks wrong. Ten successful runs show that the system can perform the task under those conditions. They do not show that it can handle variation.
Production introduces variation at every boundary. Users provide ambiguous or adversarial input. External documents contain irrelevant instructions or stale data. APIs return partial results, rate-limit responses, or malformed payloads. A tool that worked in development changes its schema. A conversation grows until the most important instruction is competing for attention with thousands of low-value tokens. A retry repeats a side effect instead of safely repeating a read.
Agent loops multiply these problems. A single incorrect model decision can cause a bad tool call; the tool result then becomes evidence for the next decision; the next decision can compound the original error. If the loop has no explicit stopping condition, a temporary failure can become an expensive sequence of retries. Anthropic describes agents as LLMs using tools in a loop and specifically recommends stopping conditions, sandboxing, testing, and guardrails because autonomy brings higher cost and compounding errors.[1]
A production system therefore needs to answer questions a demo can avoid:
What state is authoritative if the process restarts?
Which tools may be called in this state, with which arguments and permissions?
How many model and tool steps are allowed?
Which failures are retryable, and which require escalation?
How is a side effect made idempotent?
How do we know whether the agent achieved the outcome rather than merely produced plausible text?
Can an engineer reconstruct what happened from logs, traces, and stored inputs and outputs?
These are ordinary software engineering questions. The difference is that the component making some of the decisions is probabilistic.
Why better prompting is not enough
Prompting matters. Clear instructions, useful examples, and precise tool descriptions improve behavior. But a prompt is guidance, not an enforcement mechanism.
A model may interpret an instruction differently after the context changes. It may generate a syntactically valid but semantically wrong argument. It may follow an instruction embedded in retrieved content. It may decide that a partial result is good enough. Even a strong model can make a mistake because the task is ambiguous, the evidence is incomplete, or the next action is not represented clearly in the available context.
This is why structured outputs, schema validation, tool permissions, timeouts, and human approval exist. They do not assume the model will always choose correctly. They make certain classes of incorrect behavior impossible, detectable, or recoverable.
OpenAI’s agent safety guidance recommends structured outputs to constrain data flow, guardrails for user input, tool approvals for sensitive operations, and trace grading and evaluations to identify mistakes.[2] Anthropic’s guidance on tools makes the same broader point from another angle: tool interfaces should be explicit, ergonomic, validated, and evaluated with realistic tasks rather than improved through intuition alone.[3]
A useful rule is:
Put preferences in prompts. Put invariants in code.
“Prefer concise answers” belongs in a prompt. “Never issue a refund above €500 without approval” belongs in a policy check that runs independently of the model.
LLM, agent, and production AI system
These terms are often used as if they describe the same thing, but they refer to different layers.
An LLM maps input tokens to output tokens. It may return prose, structured data, or a tool request. It has no inherent authority to send an email, modify a database, or charge a customer. Those capabilities come from the application around it.
An AI agent is an application pattern in which a model helps decide what to do next, often by selecting tools and incorporating the resulting observations into a loop. Anthropic distinguishes predictable workflows, where code defines the path, from agents, where the model dynamically directs its process and tool use.
A production AI system contains the model and agent loop plus everything required to operate safely and repeatably: persistent state, authentication, tool adapters, policy enforcement, validation, retry and timeout behavior, observability, evaluation, cost controls, deployment procedures, and human escalation.
The model may be the most sophisticated component, but it is only one component.
What harness engineering means
The word “harness” is useful because it describes a control system rather than a single framework. A harness gives the model room to perform useful work while defining the boundaries of that work.
The harness owns questions such as:
What goal and state are the model allowed to see?
What actions are available right now?
Which outputs are acceptable?
What happens after success, failure, timeout, or ambiguity?
When must the system stop and ask a person?
OpenAI’s 2026 account of using Codex to build a product makes this concrete. The engineering work shifted toward repository structure, documentation, observability, tests, feedback loops, and environments that made agent behavior legible and enforceable.[4] Google’s 2026 discussion of harness engineering similarly emphasizes behavioral evaluations that check observable actions such as whether an agent used a validator or asked for clarification, rather than relying only on a single end-to-end score.[5]
Harness engineering is not a claim that autonomy should disappear. It is a recognition that autonomy must run inside an engineered operating envelope.
The components around the model
State management
State should be explicit, durable, and separate from the model’s transient context. A production request may need a conversation ID, workflow state, tool results, approval status, retry count, budget usage, and an idempotency key.
Do not make the prompt transcript the only database. If the worker restarts after a successful payment API call but before it records completion, the system needs a durable way to determine whether the action should be retried. For scalable services, Google recommends a stateless agent application with session state stored externally so any instance can resume a request.[6]
Tool constraints
Tools are capabilities, not suggestions. Give an agent the smallest useful surface area, with narrow schemas and explicit permissions. Separate read tools from write tools. Use typed arguments, allowlists, authentication at the tool boundary, and idempotency for side effects.
The model can request refund_order(order_id, amount). It should not receive a general-purpose database connection and be trusted to remember the organization’s policy. The application should validate the amount, check the user’s authorization, and require approval when the operation is consequential.
Context management
Context is a finite operational resource. More tokens do not automatically mean more understanding. Anthropic describes context as having diminishing returns and recommends keeping it informative and tight, retrieving information just in time, compacting long histories, and maintaining structured notes outside the context window.[7]
A harness can summarize old turns, retain decisions and unresolved issues, clear redundant tool output, retrieve records by identifier, and limit tool responses through pagination or filtering. The objective is not to give the model everything. It is to give it the smallest high-signal context needed for the next decision.
Validation and guardrails
Validation should happen before and after model use. Before the call, validate user input, identity, permissions, and available budget. After the call, validate the output schema, allowed transitions, referenced entities, and business rules.
Guardrails can redact sensitive input, detect suspicious instructions, block disallowed actions, and verify that retrieved data is not being passed blindly into a privileged tool. They reduce risk; they do not make an agent infallible. For high-impact actions, a human approval step is often the right boundary.
Retries and timeouts
Every external call needs a timeout. Every retry needs a policy. Retry transient network failures with bounded exponential backoff. Do not automatically retry validation failures, authorization failures, or a write whose completion status is unknown unless the operation is idempotent and the system can safely reconcile it.
The model loop also needs limits: maximum turns, maximum tool calls, maximum wall-clock time, and maximum token or monetary budget. A loop that has not reached a valid terminal state by the limit should produce a controlled failure, not continue invisibly.
Observability
Log the run ID, user or tenant scope, model version, prompt or policy version, state transitions, tool calls, validation results, latency, token usage, errors, and final outcome. Capture enough information to reproduce the decision without logging secrets or unnecessary personal data.
A trace should show the whole path: model call, selected tool, tool arguments, tool result, guardrail decision, retry, handoff, and completion. OpenAI’s agent evaluation guidance recommends starting with traces because they expose workflow-level issues such as wrong tool choice, missing handoffs, and policy violations.[8]
Evaluation
Evaluation is not a final benchmark performed once before launch. It is the test suite for the harness.
Begin with real failure cases. Turn “the agent sometimes forgets to verify the order” into a behavioral test that checks for verification before a write. Turn “the agent loops on a 404” into a test that checks bounded recovery. Keep end-to-end evaluations for overall outcomes, but add smaller deterministic checks for intermediate behavior. Google recommends this combination: behavioral evaluations help explain regressions, while larger evaluations verify the final destination.
Track quality alongside runtime, tool errors, number of calls, tokens, and cost. A response that is accurate but ten times more expensive may not be a production improvement.
Probabilistic intelligence needs deterministic constraints
A model samples from a distribution. Production software operates under contracts.
That mismatch does not mean a model must be reduced to a static classifier. It means the model should make the decisions that benefit from language understanding, while deterministic code controls the consequences of those decisions.
For example, the model may classify a support request, extract a customer ID, propose a response, or choose whether more information is needed. The harness should decide whether the extracted ID is valid, whether the support state permits a database lookup, whether a response contains prohibited data, and whether sending it requires approval.
A finite state machine is one possible mechanism. It makes the control surface explicit without removing model intelligence:

The LLM can propose a classification or draft. It cannot invent a transition from Received directly to Complete if the state machine does not allow it. It cannot call a write tool while the state is NeedInfo. It cannot bypass Approval because the transition is enforced outside the prompt.
This is the broader version of the LLM-plus-FSM pattern: deterministic states are not the entire architecture, but they are a useful control layer for operations that must be observable, bounded, and auditable.
A practical production architecture
A reliable agent service can be organized into six layers:
Interface layer: authenticates the request, creates a run ID, and applies request limits.
State layer: loads durable session and workflow state, including prior tool results and approval status.
Orchestrator: selects the current state, builds the minimal context, invokes the model, and enforces step, time, and budget limits.
Policy and tool layer: validates structured outputs, checks permissions, executes narrow tools, and records idempotency keys.
Verification layer: checks business rules, evidence, and completion criteria before marking the run successful.
Operations layer: emits traces and metrics, stores evaluation artifacts, and routes failures to retry, fallback, or human review.
A compact Python sketch makes the separation visible:
from dataclasses import dataclass
from enum import Enum
class State(Enum):
RECEIVED = "received"
RETRIEVE = "retrieve"
DRAFT = "draft"
APPROVAL = "approval"
COMPLETE = "complete"
FAILED = "failed"
@dataclass
class Run:
state: State
steps: int = 0
tokens: int = 0
max_steps: int = 8
max_tokens: int = 12_000
def can_transition(current: State, event: str) -> State:
allowed = {
(State.RECEIVED, "valid_request"): State.RETRIEVE,
(State.RETRIEVE, "grounded"): State.DRAFT,
(State.DRAFT, "needs_approval"): State.APPROVAL,
(State.DRAFT, "safe_response"): State.COMPLETE,
(State.APPROVAL, "approved"): State.COMPLETE,
}
try:
return allowed[(current, event)]
except KeyError as exc:
raise ValueError(f"Invalid transition: {current.value} + {event}") from exc
def enforce_budget(run: Run, tokens_used: int) -> None:
run.steps += 1
run.tokens += tokens_used
if run.steps > run.max_steps or run.tokens > run.max_tokens:
raise RuntimeError("Agent budget exceeded; stop and escalate")The model can still be used inside this loop for classification, retrieval planning, or drafting. The important point is that it does not own the entire loop. The harness owns the budget, transition table, and failure semantics.
Uncontrolled agent versus harnessed agent
Consider a customer-support agent asked to cancel an order and issue a refund.
An uncontrolled implementation might give the model a cancel_and_refund tool and ask it to “help the customer.” The model could infer the wrong order, issue a refund before checking eligibility, retry a timed-out request, or treat a malicious note in the order history as an instruction. The demo may look excellent because the happy path is short.
A harnessed implementation separates the work. It authenticates the customer, verifies the order ID, checks cancellation and refund policy, retrieves the order through a read-only tool, and asks the model to explain the options. If a write is appropriate, the harness validates the exact amount, creates an idempotency key, requires approval above a threshold, calls the refund API with a timeout, and reconciles the result if the response is ambiguous. Every transition is traced. A failed run ends in a known state with a human-readable reason.
The model is still valuable. It interprets the request and produces a useful explanation. But it is not the authority that decides whether a financial side effect is permitted.
The relationship with traditional software engineering
Harness engineering is not a replacement for software engineering. It is software engineering applied to a system whose behavior includes statistical variation.
The familiar practices still matter: clear interfaces, least privilege, contract tests, state persistence, idempotency, feature flags, incident response, change review, load testing, and rollback. What changes is the testing surface. Engineers must test not only functions and APIs, but also tool selection, recovery behavior, context construction, model upgrades, and the interaction between policy and generated output.
The model makes the system more flexible. The harness makes that flexibility operable.
What to focus on when moving a prototype to production
Do not begin by asking which agent framework to adopt. Begin by collecting the failures your prototype currently hides.
Write down the valid states and terminal outcomes. Define what “done” means in a way that can be checked. Inventory every tool and reduce permissions to the smallest useful scope. Externalize state. Add timeouts, bounded retries, and budgets before increasing autonomy. Build traces before tuning prompts. Turn real incidents into behavioral evaluations. Add human approval wherever an error would create an unacceptable consequence.
Then improve the model and prompts inside that envelope. A stronger model can reduce the number of failures, but it does not remove the need to handle the failures that remain.
The most important mindset shift is to stop treating the model output as the product. The product is the complete behavior of the system under normal conditions, degraded conditions, adversarial input, and partial failure.
Key Takeaways
The model is not the system. An LLM supplies probabilistic intelligence; the production system supplies predictable behavior.
Prompting is guidance, not enforcement. Put invariants, permissions, budgets, and side-effect controls in code.
Agents need a harness. State management, constrained tools, context management, validation, retries, timeouts, guardrails, observability, evaluations, and human approval are system components.
Autonomy must be bounded. Use explicit terminal states, maximum steps, time limits, cost limits, and recovery paths.
FSMs are one useful control mechanism. A deterministic state machine can constrain transitions while the model handles interpretation and drafting.
Evaluate behavior, not only final prose. Trace tool calls, state changes, policy decisions, and failure handling.
Production engineering starts with the 11th run. Reliability is defined by what happens when the happy path stops being happy.
References
[1]: https://www.anthropic.com/engineering/building-effective-agents "Building effective agents — Anthropic, December 2024"
[2]: https://developers.openai.com/api/docs/guides/agent-builder-safety "Safety in building agents — OpenAI Developer Documentation, accessed October 2026"
[3]: https://www.anthropic.com/engineering/writing-tools-for-agents "Writing effective tools for AI agents — Anthropic, September 2025"
[4]: https://openai.com/index/harness-engineering/ "Harness engineering: leveraging Codex in an agent-first world — OpenAI, February 2026"
[5]: https://developers.googleblog.com/the-anatomy-of-harness-engineering-how-to-evaluate-iterate-and-guard-ai-coding-agents/ "The Anatomy of Harness Engineering — Google for Developers, September 2026"
[6]: https://docs.cloud.google.com/architecture/choose-agentic-ai-architecture-components "Choose your agentic AI architecture components — Google Cloud Architecture Center, April 2026"
[7]: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents "Effective context engineering for AI agents — Anthropic, September 2025"
[8]: https://developers.openai.com/api/docs/guides/agent-evals "Evaluate agent workflows — OpenAI Developer Documentation, accessed October 2026"