What Is Terminal-Bench? How to Read the Score
What is Terminal-Bench? It tests whether AI agents can finish real tasks in a command line. Learn what a 4.0 score means and how to read a quoted number.
Terminal-Bench is a benchmark that tests whether an AI agent can complete real software tasks by working in a command line, with no graphical interface. Each task comes with its own tests, and the agent passes only if it satisfies them. The score you see in a launch post is the share of tasks passed. It has become a headline number in model launches, including both Claude Sonnet 5.5 and Claude Opus 5.5 this month.
What Terminal-Bench actually measures
The agent gets a terminal and a task, then has to get from the starting state to a working result: writing and debugging code, setting up an environment, recovering when a command fails. Leaderboard write-ups such as Vals AI's Terminal-Bench 4.0 page describe a task as passing only if every one of its tests passes. There is no partial credit for a mostly working answer.
That makes it a test of the whole agent, not just the model. The model plus the scaffolding around it (how it reads output, how it retries, how long it may run) produces the result. Two labs running the same model in different harnesses can post different scores.
What changed in version 4.0
The maintainers describe 4.0 on tbench.ai as a maintenance release of 3.0, not a new test. The reported changes:
8 tasks removed and 19 fixed
Task resources (time, CPU, memory) recalibrated
A uniform 8-hour agent timeout on every task, which frontier models rarely reach
Saturated tasks removed, meaning ones where every model class in the latest generation solved them every time
The project is described as a continuous benchmark hosted by Stanford, Harbor and the Laude Institute, with community contributions. The authors also note large variance in agent execution time for some models and room to improve on both cost and performance.
Three rules for reading a Terminal-Bench number
Check the version. A 3.0 score and a 4.0 score are not on the same test, because tasks were removed and fixed. If a launch table mixes them, the comparison is broken.
Check the harness. Ask who ran it, with what agent scaffold, and how many attempts. A vendor-run number and an independent leaderboard number are different kinds of evidence.
Check the size of the jump. Recent launch coverage reported a Terminal-Bench 4.0 score of 70.6% for Sonnet 5.5 against 10.3% for Sonnet 5. A gap that large between neighbouring models is a prompt to look for what else changed, not just a sign of a better model. Our note on how to tell when a benchmark score is misleading lists the usual causes.
What a good score does not tell you
Terminal-Bench measures terminal tasks with automated checks. It does not tell you how a model handles your codebase, your conventions or your ambiguous tickets. A high score is a reason to include a model in your own test, not a reason to skip the test. Our guide to benchmarking an AI coding agent on your own codebase shows how to build that test.
For the general idea behind all of this, start with what an AI benchmark is, and for the ranking sites that repeat these numbers, what an AI benchmark leaderboard is.
A worked reading of a launch table
Say a launch post lists three models with Terminal-Bench 4.0 scores of 50%, 62% and 70%. Before ranking them, ask four questions. Did every row come from the same harness and the same effort setting? Is the 70% from the vendor's own run or an independent one? How many attempts were allowed per task? And what is the noise: with a benchmark of a few dozen tasks, one task is worth a couple of percentage points, so a 2-point gap can be a single task flipping. If you cannot answer those from the post, treat the table as marketing and the ranking as unproven. If you can, it is a reasonable shortlist filter, and the next step is your own test.
FAQ
What is Terminal-Bench 4.0?
It is the current version of a benchmark that tests AI agents on real terminal tasks, released as a maintenance update to 3.0 with tasks fixed, removed and recalibrated.
How is Terminal-Bench scored?
Each task has its own tests and passes only if all of them pass. The score is the percentage of tasks passed.
Can I compare Terminal-Bench 3.0 and 4.0 scores?
Not directly. The tasks and resource limits changed between versions, so the numbers come from different tests.
Who runs Terminal-Bench?
It is described as a continuous benchmark hosted by Stanford, Harbor and the Laude Institute, with community contributions on GitHub.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


