Benchmarks & Datasets

Standard tasks and datasets for comparing agents. Useful for direction, but read what a benchmark actually measures before treating a leaderboard position as a selection criterion.

10 projects

promptfooObservability

Test cases are YAML: a prompt, some inputs, and assertions about the output. That is enough to produce a side-by-side matrix comparing several prompts or models on the same inputs, and to fail a build when a change makes things worse. The same tool also attacks the application, generating adversarial inputs to surface vulnerabilities and compliance risks. It works from the outside — prompt in, answer out — so when the question is which step inside a running agent went wrong, this is not the instrument for it.

OfficialMIT24.6k+208
DeepEvalObservability

The shape is deliberately pytest's: assertions inside test functions, a runner that discovers them, parameterised cases — except the assertion is about whether an answer was faithful to its context or whether an agent finished the task. That framing is the point, because it puts quality regressions in front of the same gate that catches syntax errors. The thing to plan for is that most of these metrics are scored by a model, so the suite has a bill and a variance that ordinary tests do not.

OfficialApache-2.017.9k+174
camelOrchestration

CAMEL's core idea is agents talking to each other: a RolePlaying session pairs an AI user that issues instructions with an AI assistant that carries them out, and a Workforce hands split-up subtasks to workers under a coordinator. Around that sit eighty-odd toolkits, around fifty model backends, benchmark harnesses such as GAIA and BrowseComp, and pipelines that turn the resulting conversations into fine-tuning data. That breadth is the point if you are running experiments; if you only want one dependable agent in production you will carry a great deal you never use, and the library has been public since 2023 yet is still on 0.2.x.

OfficialApache-2.017.7k+35
garakGuardrails

Point it at a chatbot or a model and it works through families of attacks — prompt injection, jailbreaks, encoding-based bypasses, data leaks and replay, toxicity, false reasoning — and reports which ones got through. Keeping the attacks, the judgement and the connection to the target as separate components is what lets a new probe or a new backend be added without rewriting the rest. It is the finding half of security, not the fixing half: the output tells you which probes succeeded, not what policy to write in response.

OfficialApache-2.09.1k+177
SWE-benchBenchmarks

A task gives the model a codebase and an issue and asks for a patch; the patch is accepted only if the tests that were failing now pass and the ones that were passing still do. Using the projects' own test suites as the judge is what makes the result mean something concrete rather than resemble a rubric score. What it measures is correspondingly narrow: producing a fix for a problem someone else already found, filed and described — not deciding what to build, not reviewing, not navigating a codebase without an issue to anchor on.

OfficialMIT5.7k+51
AgentBenchBenchmarks

A model that answers well can still be poor at operating something, and a question-and-answer benchmark will not show that. AgentBench puts a model into environments that push back: a shell it has to work in, a database it has to query, a knowledge graph, a text world and a shopping site where the task only counts if the right item ends up ordered. Every task is scored on whether it was completed, not on whether the reasoning read well. The current release runs those environments as containers driven by function calling, which makes it reproducible on your own hardware — with the practical caveats the maintainers list, including one environment that needs about 16GB of memory and another that leaks until its worker is restarted.

Apache-2.03.7k
OSWorldBenchmarks

It is the common yardstick the computer-use projects cite, and the reason it carries weight is that the environment is genuine: Ubuntu or Windows in a VM, running real applications, reachable through VMware, VirtualBox, Docker with KVM, or AWS, Azure and Modal for parallel runs. That realism is also the cost — adopting it means hypervisors, VM images and cleanup after interrupted runs. Results are version-sensitive too: the Verified update changed tasks and signals, and the project asks that comparisons be made against the re-run numbers.

OfficialApache-2.03.1k+11
Inspect AIBenchmarks

Where the other entries in this category are each one measurement, this is the machinery for building measurements: a dataset supplies inputs and targets, a solver produces a response (a single model call or a full multi-turn agent), a scorer decides whether it was right, and a task binds the three together. More than two hundred pre-built evaluations come with it, model-written code runs inside a sandbox, and a log viewer shows what actually happened on each sample. What it cannot do is tell you what is worth measuring.

OfficialMIT2.7k+59
Windows Agent ArenaBenchmarks

Almost every desktop-agent benchmark runs on Linux, and almost every desktop that matters commercially runs Windows. This is Microsoft's answer to that gap — a reproducible Windows environment with tasks across the applications people actually use, scored on whether the task was completed. Its useful trick is parallelism: runs can be spread across cloud machines so a few hundred tasks finish in minutes instead of over a day, which is what makes it usable during development rather than only at publication time. Read the last commit date before planning around it, though — this is a research artefact tied to a 2024 paper, and it has moved slowly since.

OfficialMIT891
Terminal-BenchBenchmarks

Rather than a fixed set that saturates and stops distinguishing anything, this is a continuous benchmark: tasks are added and fixed over time and the dataset is published as tagged releases, run by a separate harness called Harbor. That design is the strength and the complication — a number without a dataset version attached does not mean much, version 1.0 lives in its own repository that the original URL now redirects to, and the documented first step is running the reference solutions repeatedly to confirm your own sandbox is sound before trusting any result.

OfficialApache-2.0556+32