What Is a Long-Horizon Task in AI?
A long-horizon task is any job an AI has to carry across many dependent steps without losing the thread. It is the hardest thing agents do, and the failures are specific.
A long-horizon task is one an AI has to carry across many dependent steps, where step 40 depends on a decision made at step 3 and nobody reminded it. Migrating a database schema, working through a bug that spans four files, running a research thread over an afternoon. The opposite is a short-horizon task: summarise this, translate that, one question and one answer. Models have been good at short-horizon work for years. Long-horizon work is where the current generation is still visibly imperfect, and it is what almost every 2026 model release is now benchmarked on.
Why length is the hard part
The intuition that a long task is just many short tasks is wrong in a specific way. Errors compound. If a model gets each step right 98% of the time and steps are independent, 40 steps finish clean about 45% of the time. At 95% per step, it is 13%. Nothing is broken in that model. The per-step accuracy is excellent. The task still fails more often than not.
This is why a jump from 95% to 98% per-step reliability feels enormous in agent work and barely registers in chat. It is also why benchmark gains on long-horizon suites tend to be dramatic when they come at all: you are not measuring capability, you are measuring the product of many capabilities.
The four ways long-horizon tasks actually fail
Goal drift
The agent optimises for the most recent instruction rather than the original one. You asked it to fix a failing test without changing behaviour; twenty steps later it is refactoring the module because that was the last thing you discussed. The original constraint is still in the context, it is just no longer weighted heavily.
Context exhaustion
Long tasks generate output, and output becomes input. Eventually early context gets truncated or summarised away, taking the requirements with it. This is the failure mode behind most agents that start strong and drift into nonsense, and it is why understanding what a context window actually holds is a practical question rather than a trivia one.
Unverified intermediate state
The agent assumes step 12 worked because it looked like it worked, then builds nine more steps on top. Human engineers check; agents often do not, unless the harness forces a check. This is the single highest-leverage thing you can fix from outside the model.
No stopping rule
The task has no clear completion test, so the agent either stops early on something plausible or loops on refinements nobody asked for. Long-horizon work needs an explicit finish condition more than short work does, because there is more room to wander.
How benchmarks measure it
Two families of benchmark have become the reference points for this in 2026, and both appeared in releases published this week.
Benchmark | What it runs | Recent published figure |
|---|---|---|
Terminal-Bench 4.0 | Multi-step work in a real shell | Claude Fable 5.1 at 55.8%, up from 42.0% for Fable 5 |
Terminal-Bench-Science 0.1 | Scientific workflows end to end | Claude Fable 5.1 at 52.6%, up from 24.7% |
DeepSWE v1.1 | Long-horizon software engineering | Google says Gemini 3.8 Flash beats most larger frontier models |
OSWorld 2.0 (strict) | Computer use across applications | Claude Fable 5.1 at 41.7%, up from 36.1% |
The pattern is worth noticing: absolute scores sit in the 40% to 56% range on tasks a competent human finishes almost every time. Long-horizon agent work in September 2026 is genuinely useful and genuinely unreliable at the same time, and the benchmark numbers say so plainly. Read them alongside what an AI benchmark actually measures before drawing conclusions.
What this changes about how you use agents
Shorten the horizon deliberately. Three supervised five-step runs beat one unsupervised fifteen-step run, because you catch the compounding error while it is still cheap.
Make state checkable. Give the agent a command that tells it whether the last step actually worked, and require it to run it. Tests, a type check, a health endpoint, anything cheap and unambiguous.
Restate the goal at checkpoints. Not because the model has forgotten, but because restating re-weights it against everything that has accumulated since.
Write the stopping condition before you start. If you cannot state what done looks like in one sentence, the agent has no chance of recognising it either.
On the practical side of that, see how long to let an AI coding agent run unattended, which puts numbers on the supervision tradeoff, and what agentic AI means for the broader category. The underlying mechanics live in how AI models work.
Frequently asked questions
Is a long-horizon task the same as a long context?
No. Long context is about how much text the model can hold at once. Long horizon is about how many dependent decisions it has to make in sequence. You can have a long-horizon task with a small context, and a long context with a single-step question.
How many steps counts as long-horizon?
There is no fixed threshold, but in practice the failure modes above start showing up somewhere past ten dependent steps, and get sharply worse past thirty.
Does a bigger model fix this?
Partly. Bigger models tend to have higher per-step reliability, which compounds favourably. But the September 2026 benchmark scores show even frontier models finishing barely half of these tasks, so scale has not solved it.
Why do agents do better with tests in the repository?
Because a test suite is a cheap, unambiguous check on intermediate state, which removes the third failure mode above. It is the closest thing to a fix that does not require a better model.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


