Atlas / BUILD / Agents / Long-Running Agents
DEEP-DIVE · AGENTS

Long-Running Agents: From Minutes to Days

Agents now sustain work measured in hours, and some products ship agents that never switch off. What breaks when a task gets long, how durability, memory and supervision change the architecture, and where always-on agents fit in an enterprise.

TL;DR
  • The length of task frontier agents can complete is growing fast. METR's time-horizon measure doubled roughly every seven months over its full history and closer to every four months since 2024, and one early-2026 frontier model was estimated at sixteen hours or more at 50% reliability. Hours-long agent work is now an engineering problem, not a research one.
  • Long tasks fail differently from short ones: context runs out, small errors compound, budgets run away, infrastructure restarts and the world changes mid-task. The fixes are architectural: durable execution with checkpoints, idempotent tools, external state and compaction, enforced budgets and asynchronous supervision.
  • Always-on agents, persistent workers with their own compute, standing goals and event triggers, reached mainstream products in 2026. Treat each as a non-human worker: its own identity, scoped permissions, an accountable human manager, a budget, an audit trail and an off switch.

The time horizon is moving

The most useful single measure of agent progress is METR's time horizon: the length of task, expressed as the time a skilled human would need, that an agent completes at a given success rate, usually 50%. Over its full history the measure doubled roughly every seven months. METR's 2026 update found that for frontier models released since 2024 the doubling time had shortened to around four months. In March 2026 it estimated an early version of one frontier model at a 50% horizon of at least sixteen hours, with a wide confidence interval and at the upper edge of what its task suite can measure.

Read those numbers with care. A 50% success rate is not production reliability, and the horizon at higher reliability levels is several times shorter. The tasks are weighted toward software engineering and research, and your workloads are not METR's suite. But the direction is unambiguous, and products have followed it: coding agents that work for hours in cloud sandboxes, research agents that assemble reports across long sessions, and in September 2026 a major vendor's always-on agents, each with its own cloud computer and browser, that take a goal, connect to the apps they need, and return results for review.

The architectural consequence is that a design built for a thirty-second chat turn does not stretch to an eight-hour task. It does not merely take longer. It fails in different ways.

What breaks when the task gets long

Most long-task failures are short-task failures with time to compound. A step that succeeds 98% of the time looks reliable in a demo; chained a hundred times without verification, the whole run succeeds about 13% of the time. The failure modes worth designing for:

FailureWhy length makes it worseMitigation
Context exhaustionThe transcript outgrows the window; early instructions degrade or drop outCompaction, external notes, guides re-injected on a cadence
Compounding errorAn early mistake propagates through every later stepVerified checkpoints; sensor gates between steps
Goal driftThe agent slowly loses the original goal or a constraintA plan file the agent re-reads; periodic checks against the spec
Cost runawayLoops and retries accumulate silently over hoursPer-run budgets for tokens, money and wall-clock time, enforced as hard stops
Infrastructure failureDeploys, restarts, timeouts and rate limits all happen eventuallyDurable execution, resumable steps, idempotent tools
Stale worldData or system state changes while the task runsRe-read before write; leases and optimistic concurrency
Absent humansAn approval is needed when nobody is watchingAsynchronous approvals that pause the run in a safe state

None of these is exotic. Distributed-systems engineers have handled versions of every row for decades. What is new is that the component deciding the next step is a probabilistic model, which makes verification between steps mandatory rather than optional.

Durability: agents as workflows

The most productive mental model is that a long-running agent is a long-running workflow whose next step happens to be chosen by a model. Workflows that run for hours are built on durable execution: every step's input and output is persisted, a crash resumes from the last completed step, and retries are safe. Engines such as Temporal, Restate and DBOS, cloud workflow services, and the checkpointers built into agent frameworks all provide some version of this. Choose one deliberately; do not let each team improvise resumption logic.

   goal + budget + scope lease
             |
             v
         [ PLAN ] <-------------------------------+
             |                                    |
             v                                    |
       [ STEP n ] --> tool call (idempotency key) |
             |                                    |
             v                                    |
      [ SENSORS ] -- fail --> repair / retry -----+
             |
           pass
             |
             v
      [ CHECKPOINT ]  persist state + evidence
             |
     +-------+---------------+------------------+
     |                       |                  |
 next step          budget hit or         goal met
 (back to PLAN)     approval needed           |
                             |                v
                             v        deliver with evidence
                  pause safely + notify
A long-running agent as a durable loop: each verified step is checkpointed, tools are idempotent, and budgets or approvals pause the run rather than ending it.

Two engineering rules make durability work with models in the loop. First, tools must be idempotent, because resumption replays steps. A payment, an email or a ticket creation needs an idempotency key so that replaying the step after a crash does not do it twice. Second, record model outputs as part of the step's persisted result. A replay should reuse the decision the model already made, not ask again and get a different answer, which is the same discipline workflow engines apply to any non-deterministic call.

Durability also changes how you think about plans. A plan written at the start of an eight-hour task is a hypothesis. Persist it, let the agent revise it at checkpoints, and keep the revision history, because the path from the original plan to the delivered result is the best evidence an auditor or a debugging engineer will get. The decomposition patterns in the planning deep-dive apply, with the addition that every node of the plan must be resumable.

Context and memory over hours

The context window is working memory, not storage, and long runs exhaust it. Three techniques keep a long task coherent. Compaction summarizes older turns while preserving the decisions made, the constraints in force and the open questions, so the agent keeps the substance without the transcript. External state moves the important things out of the window entirely: a plan file, a task list with statuses, a decisions log and a progress note the agent reads at each step and updates as it goes. Re-injection restates the goal and the critical guides at intervals, because instructions buried deep in a long context get less attention than those near the end.

Structured state beats a transcript. A task list with twelve items marked done, in progress or blocked is easier to resume, hand off and audit than four hundred messages. It also lets a second agent, or a human, pick up the work, which matters for runs that outlast a working day. The memory architectures in the agent memory deep-dive and the budget discipline in the context engineering playbook carry over directly; long runs simply make their absence fatal instead of merely expensive.

Cost deserves its own line. A long run re-sends a large, mostly stable prefix many times, which is exactly the pattern prompt caching rewards. Keep the stable parts of the context, such as guides and tool definitions, at the front and unchanged, and let the volatile parts change at the end.

Supervision without babysitting

Nobody can watch an eight-hour run, and nobody should have to. Supervision for long-running agents is designed in advance and enforced by the harness, not by a person staring at a log. Six mechanisms cover most needs:

Budgets are features. A run that stops at its budget and explains why is a working system. A run that quietly spends ten times its plan is an incident, whether or not its output was good. Put the budget in the harness, never only in the prompt, because a model can talk itself past a sentence but not past a hard limit.

Always-on agents

A long-running agent works on one task for a long time. An always-on agent is a different shape: not a task but a role. It has a persistent identity, standing goals, its own compute such as a cloud desktop and browser, and triggers that wake it: an email arrives, a ticket is opened, a schedule fires, a metric crosses a threshold. It accumulates memory of preferences and context across weeks. In 2026 these moved from prototypes into mainstream products, both as personal agents inside chat assistants and as AI coworkers that data and application platforms run on top of a shared enterprise context layer.

For an enterprise, the right frame is that an always-on agent is a non-human worker, and every control you apply to workers applies to it. It needs its own identity rather than a borrowed human session, as the identity and tenancy deep-dive argues. It needs scoped permissions that are reviewed periodically, because standing agents accumulate access the way long-tenured employees do. It needs a budget, an audit trail, and a decommissioning process for when the role ends. And it needs to appear in an inventory, because an organization that cannot list its always-on agents cannot govern them.

The security surface is the part most often underestimated. An agent that reads an inbox treats every incoming message as input, which makes every sender a potential source of prompt injection. Standing access plus untrusted input plus the ability to act is the combination that turns a helpful assistant into an exfiltration channel. Keep the dangerous capabilities apart, gate outbound actions, and log everything.

An always-on agent needs a manager. Name the human accountable for what each one does, review its work on a cadence the way you would a new hire's, and give that person the off switch. Accountability that is spread across "the platform team" belongs to nobody.

The architect view

Long-running agents move the bottleneck from model capability to systems engineering. The models can now sustain the work; whether an enterprise can let them depends on durability, state management, budgets and supervision, which are platform concerns with well-understood solutions from the workflow and distributed-systems world.

Four commitments follow. Offer durable execution as a platform service, so no team hand-rolls resumption logic. Enforce budgets and scope leases in the harness, never only in prompts. Standardize structured state and progress reporting, so every long run can be resumed, handed off and audited. And govern always-on agents as part of the workforce: identity, a named manager, an inventory entry and a decommissioning path.

Start with long tasks that have strong sensors: code migrations with test suites, reconciliations with exact expected totals, research syntheses with checkable citations, test generation against coverage targets. Measure your own time horizon, the longest task class where completion and intervention rates meet your bar, and re-measure with every model generation, because it will move. The question is no longer whether an agent can work through an afternoon. It is whether your systems let it do so safely, resumably and on budget.

← Agent Skills: Packaging Know-How for Agents ALL OF AGENTS