
promptfoo
Declarative evals and automated red-teaming from one CLI, running entirely on your own machine
What is promptfoo?
Test cases are YAML: a prompt, some inputs, and assertions about the output. That is enough to produce a side-by-side matrix comparing several prompts or models on the same inputs, and to fail a build when a change makes things worse. The same tool also attacks the application, generating adversarial inputs to surface vulnerabilities and compliance risks. It works from the outside — prompt in, answer out — so when the question is which step inside a running agent went wrong, this is not the instrument for it.
What can you do with promptfoo?
- Describe the test, not the harness — A YAML file with prompts, inputs and assertions is the whole configuration — no notebook to maintain and no test framework to wire up first, which also keeps it usable from a codebase in any language.
- Compare options as a table — Results render as a matrix across prompts and inputs in a web viewer and on the command line, so the choice between two prompts or two models is settled by a grid rather than by impressions.
- Fail the build on a regression — It runs as a CLI, as a library or inside CI pipelines such as GitHub Actions, which is what turns evaluation from an occasional exercise into a gate.
- Attack your own application first — Automated red-teaming generates adversarial inputs to find vulnerabilities and compliance risks, covering the security half that evaluation tools in this category usually leave out.
- Keep the data on your machine — Evaluations run locally and talk to the model provider directly, so test inputs and outputs are not routed through a third-party service to be scored.
Before you choose promptfoo
- It exercises the application from the outside, so it does not replace tracing: when the failure is inside a multi-step agent, this tells you the answer got worse but not which step caused it.
Star history
21 Aug to 28 Aug · +208
Frequently asked questions
Is promptfoo free for commercial use?
promptfoo is released under the MIT licence — OSI-approved open source, which permits commercial use.
How can promptfoo be deployed?
promptfoo is available as Runs locally / Self-hosted.
Documentation
Reproduced from the promptfoo/promptfoo README, published under MIT. Read the original ↗
Promptfoo: LLM evals & red teaming
Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed. Read the company update.
Quick Start
Requires Node.js >=22.22.0 for npm and npx usage. Node.js 24 LTS
is recommended; see the runtime support guide.
npm install -g promptfoo
promptfoo init --example getting-started
Also available via brew install promptfoo and pip install promptfoo. You can also use npx promptfoo@latest to run any command without installing.
Most LLM providers require an API key. Set yours as an environment variable:
export OPENAI_API_KEY=sk-abc123
Once you’re in the example directory, run an eval and view results:
cd getting-started
promptfoo eval
promptfoo view
See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.
What can you do with Promptfoo?
- Test your prompts and models with automated evaluations
- Secure your LLM apps with red teaming and vulnerability scanning
- Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
- Automate checks in CI/CD
- Review pull requests for LLM-related security and compliance issues with code scanning
- Share results with your team
Here’s what it looks like in action:
It works on the command line too:
It also can generate security vulnerability reports:
Why Promptfoo?
- Developer-first: Fast, with features like live reload and caching
- Private: LLM evals run 100% locally - your prompts never leave your machine
- Flexible: Works with any LLM API or programming language
- Battle-tested: Powers LLM apps serving 10M+ users in production
- Data-driven: Make decisions based on metrics, not gut feel
- Open source: MIT licensed, with an active community