← Back to all projects

promptfoo

Declarative evals and automated red-teaming from one CLI, running entirely on your own machine

OfficialMIT
Stars
24.6k
Forks
2.2k
Open issues
537
Last commit
28 Aug 2026

What is promptfoo?

Test cases are YAML: a prompt, some inputs, and assertions about the output. That is enough to produce a side-by-side matrix comparing several prompts or models on the same inputs, and to fail a build when a change makes things worse. The same tool also attacks the application, generating adversarial inputs to surface vulnerabilities and compliance risks. It works from the outside — prompt in, answer out — so when the question is which step inside a running agent went wrong, this is not the instrument for it.

What can you do with promptfoo?

  • Describe the test, not the harness — A YAML file with prompts, inputs and assertions is the whole configuration — no notebook to maintain and no test framework to wire up first, which also keeps it usable from a codebase in any language.
  • Compare options as a table — Results render as a matrix across prompts and inputs in a web viewer and on the command line, so the choice between two prompts or two models is settled by a grid rather than by impressions.
  • Fail the build on a regression — It runs as a CLI, as a library or inside CI pipelines such as GitHub Actions, which is what turns evaluation from an occasional exercise into a gate.
  • Attack your own application first — Automated red-teaming generates adversarial inputs to find vulnerabilities and compliance risks, covering the security half that evaluation tools in this category usually leave out.
  • Keep the data on your machine — Evaluations run locally and talk to the model provider directly, so test inputs and outputs are not routed through a third-party service to be scored.

Before you choose promptfoo

  • It exercises the application from the outside, so it does not replace tracing: when the failure is inside a multi-step agent, this tells you the answer got worse but not which step caused it.

Star history

21 Aug to 28 Aug · +208

24.4k24.6k

Frequently asked questions

Is promptfoo free for commercial use?

promptfoo is released under the MIT licence — OSI-approved open source, which permits commercial use.

How can promptfoo be deployed?

promptfoo is available as Runs locally / Self-hosted.

Documentation

Reproduced from the promptfoo/promptfoo README, published under MIT. Read the original ↗

Promptfoo: LLM evals & red teaming

Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed. Read the company update.

Quick Start

Requires Node.js >=22.22.0 for npm and npx usage. Node.js 24 LTS is recommended; see the runtime support guide.

npm install -g promptfoo
promptfoo init --example getting-started

Also available via brew install promptfoo and pip install promptfoo. You can also use npx promptfoo@latest to run any command without installing.

Most LLM providers require an API key. Set yours as an environment variable:

export OPENAI_API_KEY=sk-abc123

Once you’re in the example directory, run an eval and view results:

cd getting-started
promptfoo eval
promptfoo view

See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.

What can you do with Promptfoo?

  • Test your prompts and models with automated evaluations
  • Secure your LLM apps with red teaming and vulnerability scanning
  • Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
  • Automate checks in CI/CD
  • Review pull requests for LLM-related security and compliance issues with code scanning
  • Share results with your team

Here’s what it looks like in action:

It works on the command line too:

It also can generate security vulnerability reports:

Why Promptfoo?

  • Developer-first: Fast, with features like live reload and caching
  • Private: LLM evals run 100% locally - your prompts never leave your machine
  • Flexible: Works with any LLM API or programming language
  • Battle-tested: Powers LLM apps serving 10M+ users in production
  • Data-driven: Make decisions based on metrics, not gut feel
  • Open source: MIT licensed, with an active community

Learn More