
promptfoo
宣言的な評価と自動レッドチーミングを1つのCLIで。処理はすべて手元のマシンで完結する
promptfooとは
テストケースはYAMLです。プロンプト、いくつかの入力、そして出力に対する表明を書きます。それだけで、同じ入力に対して複数のプロンプトやモデルを並べた比較表が得られ、変更で品質が下がったときにビルドを失敗させられます。同じツールがアプリケーションへの攻撃も行い、敵対的な入力を生成して脆弱性やコンプライアンス上のリスクを洗い出します。ただし外側から、プロンプトを入れて回答を見る形で動きます。稼働中のエージェントのどの段階で問題が起きたかを知りたい場合、この道具は向きません。
promptfooで何ができますか?
- 基盤ではなくテストそのものを書く — プロンプト、入力、表明を書いたYAMLファイルが設定のすべてです。保守すべきノートブックも、先に組み上げるテストフレームワークもありません。どの言語のコードベースからでも使えます。
- 選択肢を表で比べる — 結果はプロンプトと入力を軸にした行列として、Web画面とコマンドラインの双方に表示されます。2つのプロンプトやモデルの優劣を、印象ではなく表で決められます。
- 劣化したらビルドを止める — CLIとしても、ライブラリとしても、GitHub ActionsなどのCIパイプラインの中でも動きます。評価を、思い出したときの作業から通過必須の関門に変えられます。
- 先に自分のアプリを攻撃する — 自動レッドチーミングが敵対的な入力を生成し、脆弱性とコンプライアンス上のリスクを洗い出します。同分類の評価ツールが手を出さないセキュリティ側を覆います。
- データを手元に置いたままにする — 評価はローカルで実行され、モデルプロバイダとは直接やり取りします。テストの入出力が採点のために第三者のサービスを経由することがありません。
promptfooを選ぶ前に
- アプリケーションを外側から動かすため、トレースの代わりにはなりません。多段のエージェント内部で失敗した場合、回答が悪化したことは分かっても、どの段階が原因かは分かりません。
スター推移
8月21日〜8月28日 · +208
よくある質問
promptfooは商用利用できますか?
promptfooはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
promptfooはどの形で使えますか?
promptfooはローカル実行・セルフホストの形で利用できます。
ドキュメント
promptfoo/promptfoo のREADMEより転載(MIT)。 原文を読む ↗
Promptfoo: LLM evals & red teaming
Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed. Read the company update.
Quick Start
Requires Node.js >=22.22.0 for npm and npx usage. Node.js 24 LTS
is recommended; see the runtime support guide.
npm install -g promptfoo
promptfoo init --example getting-started
Also available via brew install promptfoo and pip install promptfoo. You can also use npx promptfoo@latest to run any command without installing.
Most LLM providers require an API key. Set yours as an environment variable:
export OPENAI_API_KEY=sk-abc123
Once you’re in the example directory, run an eval and view results:
cd getting-started
promptfoo eval
promptfoo view
See Getting Started (evals) or Red Teaming (vulnerability scanning) for more.
What can you do with Promptfoo?
- Test your prompts and models with automated evaluations
- Secure your LLM apps with red teaming and vulnerability scanning
- Compare models side-by-side (OpenAI, Anthropic, Azure, Bedrock, Ollama, and more)
- Automate checks in CI/CD
- Review pull requests for LLM-related security and compliance issues with code scanning
- Share results with your team
Here’s what it looks like in action:
It works on the command line too:
It also can generate security vulnerability reports:
Why Promptfoo?
- Developer-first: Fast, with features like live reload and caching
- Private: LLM evals run 100% locally - your prompts never leave your machine
- Flexible: Works with any LLM API or programming language
- Battle-tested: Powers LLM apps serving 10M+ users in production
- Data-driven: Make decisions based on metrics, not gut feel
- Open source: MIT licensed, with an active community