Inspect AI
The framework for writing and running evaluations, from the UK AI Security Institute — not a benchmark
What is Inspect AI?
Where the other entries in this category are each one measurement, this is the machinery for building measurements: a dataset supplies inputs and targets, a solver produces a response (a single model call or a full multi-turn agent), a scorer decides whether it was right, and a task binds the three together. More than two hundred pre-built evaluations come with it, model-written code runs inside a sandbox, and a log viewer shows what actually happened on each sample. What it cannot do is tell you what is worth measuring.
What can you do with Inspect AI?
- Assemble an evaluation from four parts — Datasets carry inputs and targets, solvers turn an input into a response, scorers judge it by text comparison, model grading or a scheme you write, and a task binds them — so swapping the scorer does not mean rewriting the evaluation.
- Start from two hundred existing evaluations — A large library of pre-built evaluations ships with the framework, which means the common measurements are available immediately and the custom work is confined to what is genuinely specific to you.
- Evaluate agents, not just completions — Built-in agents and multi-agent primitives are supported, along with bash, python, text-editing, web search, browsing and computer tools, plus custom and MCP tools — so the thing under test can be a full agent.
- Run untrusted model code safely — Sandboxing covers Docker, Kubernetes, Modal, Proxmox and Vagrant with an extension API for anything else, which is a prerequisite once an evaluation involves executing what the model wrote.
- Look at what actually happened — A web-based log viewer shows individual samples and runs, turning a surprising aggregate score into something you can trace back to specific transcripts rather than only report.
Before you choose Inspect AI
- It is a framework, not a verdict: it runs the evaluation you design, so deciding what to measure and what counts as passing stays with you and is the part that is actually hard.
Star history
21 Aug to 28 Aug · +59
Frequently asked questions
Is Inspect AI free for commercial use?
Inspect AI is released under the MIT licence — OSI-approved open source, which permits commercial use.
How can Inspect AI be deployed?
Inspect AI is available as Runs locally / Self-hosted.
Documentation
Reproduced from the UKGovernmentBEIS/inspect_ai README, published under MIT. Read the original ↗
Welcome to Inspect, a framework for large language model evaluations created by the UK AI Security Institute.
Inspect provides many built-in components, including facilities for prompt engineering, tool usage, multi-turn dialog, and model graded evaluations. Extensions to Inspect (e.g. to support new elicitation and scoring techniques) can be provided by other Python packages.
To get started with Inspect, please see the documentation at https://inspect.aisi.org.uk/.
Inspect also includes a collection of over 200 pre-built evaluations ready to run on any model (learn more at https://inspect.aisi.org.uk/evals/).
Coding agents: a structured index of the docs is published at https://inspect.aisi.org.uk/llms.txt. The user guide is concatenated as Markdown at https://inspect.aisi.org.uk/llms-guide.txt, and https://inspect.aisi.org.uk/llms-full.txt additionally bundles the API and CLI reference. Individual pages are available as Markdown by appending
.mdto the.htmlpath (e.g./extensions/index.html.md).
To work on development of Inspect, clone the repository and install with the -e flag and [dev] optional dependencies:
git clone https://github.com/UKGovernmentBEIS/inspect_ai.git
cd inspect_ai
pip install -e ".[dev]"
Alternatively, if you use uv, sync the development environment from the checked-in lockfile:
uv sync --extra dev
The uv workflow is supported but not required. The uv.lock file records a reproducible development resolution; project dependencies are still declared in requirements*.txt and exposed through pyproject.toml. When changing dependencies, update the appropriate requirements file and refresh the lockfile rather than relying on uv add.
Optionally install pre-commit hooks via
make hooks
Run linting, formatting, and tests via
make check
make test
When working in a uv-managed environment, prefix those commands with uv run (for example, uv run make check).
If you use VS Code, you should be sure to have installed the recommended extensions (Python, Ruff, and MyPy). Note that you’ll be prompted to install these when you open the project in VS Code.
Frontend development (TypeScript)
The web UI lives in a git submodule at src/inspect_ai/_view/ts-mono/. These steps are only needed if you plan to work on the TypeScript/React frontend — Python-only contributors can skip this entirely.
Initialize the submodule and install dependencies — see the one-time setup guide.
Documentation
To work on the Inspect documentation, install the optional [doc] dependencies with the -e flag and build the docs:
pip install -e ".[doc]"
cd docs
quarto render # or 'quarto preview'
If you intend to work on the docs iteratively, you’ll want to install the Quarto extension in VS Code.