← Back to all projects

Terminal-Bench

Measures agents in the environment terminal agents actually inhabit: a shell inside a container

OfficialApache-2.0
Stars
556
Forks
423
Open issues
219
Last commit
28 Aug 2026

What is Terminal-Bench?

Rather than a fixed set that saturates and stops distinguishing anything, this is a continuous benchmark: tasks are added and fixed over time and the dataset is published as tagged releases, run by a separate harness called Harbor. That design is the strength and the complication — a number without a dataset version attached does not mean much, version 1.0 lives in its own repository that the original URL now redirects to, and the documented first step is running the reference solutions repeatedly to confirm your own sandbox is sound before trusting any result.

What can you do with Terminal-Bench?

  • Run an agent against the tasks with one command — Harbor downloads the dataset and executes the run; flags select which agent and which model are under test, so comparing two models is a change of argument rather than a change of setup.
  • Verify your sandbox before trusting a number — The documented first step is running the reference solutions several times over to confirm every task behaves correctly in your environment — a flaky oracle means the scores you produce are measuring your infrastructure.
  • Run it at real concurrency in the cloud — Modal is supported as an execution environment with hundreds of concurrent runs, which is what makes a full pass over the task set finish in a sitting rather than overnight.
  • Track a frontier that keeps moving — Tasks are added and corrected continuously and published as tagged dataset releases, so the benchmark keeps discriminating as agents improve instead of going flat at the top.
  • Fix a task you disagree with — Task bugs are reported as issues and improvements or new tasks arrive as pull requests, with a publicly visible roadmap — the task set is editable rather than handed down.

Before you choose Terminal-Bench

  • The dataset is versioned separately from the harness and keeps changing by design, so a Terminal-Bench score is only comparable to another when both state which dataset release produced it.
  • Version 1.0 sits in a separate repository and the original project URL now redirects there rather than here, so older published results need checking against which of the two produced them.

Star history

21 Aug to 28 Aug · +32

524556

Frequently asked questions

Is Terminal-Bench free for commercial use?

Terminal-Bench is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.

How can Terminal-Bench be deployed?

Terminal-Bench is available as Runs locally / Self-hosted.

Documentation

Reproduced from the harbor-framework/terminal-bench README, published under Apache-2.0. Read the original ↗

Terminal-Bench

Terminal-Bench is a benchmark designed to measure the frontier of agent work with a diverse, difficult, high quality set of tasks that evolve over time. All frontier agent builders use Terminal-Bench to track progress and compare capabilities.

Terminal-Bench is a continuous benchmark, with tagged releases published on the Harbor Hub. Open an issue to report any task bugs and open a PR for task improvements or new tasks. Our roadmap is publicly visible.

Tasks

The latest published version of the dataset is available on the Harbor Hub.

Running the Benchmark

Install Harbor and run the oracle solutions 5x to confirm all tasks work as expected in your the sandboxing environment. We develop our tasks using Modal in our CI/CD and leaderboard experiments - if the oracle flakes on your setup, please open an issue.

uv tool install 'harbor[modal]'
uv run harbor run -d terminal-bench/terminal-bench@latest \
   -k 5  \
   --agent oracle \
   --n-concurrent 500 \
   --env modal

To test an agent and model, pass --agent and --model:

uv run harbor run -d terminal-bench/terminal-bench@latest \
   --agent claude-code \
   --model anthropic/claude-fable-5 \
   --ak reasoning_effort=max \
   --n-concurrent 100 \
   --env modal

If your agent runs encounter any problems, please open an issue.

Resources

Terminal-Bench
Ask AI
GitHub