
Ragas
検索を使う仕組みを、必要な文書を引けたかと、答えがその文書の中に収まっているかの2点で採点する
Ragasとは
検索を使う仕組みで最も多い失敗は、外から見えません。答えは読みやすいのに、引いてきた文書のどこにも根拠がない、という状態です。Ragasはそれを直接測ります。処理を2つに分け、必要な内容を含む文書を引けていたか、生成された答えがその範囲に収まっているかを別々に採点します。手持ちの文書から試験用のデータを作ることもできるため、最初の評価のために質問を100件手書きするところから始めずに済みます。ほとんどの指標自体がモデルの呼び出しである点は、あらかじめ織り込んでおくべきです。一通り走らせれば費用がかかり、同じデータで2回実行しても数値は完全には一致しません。
Ragasで何ができますか?
- 検索の失敗と生成の失敗を切り分ける — 検索と回答を別々に採点するため、直すべきなのが索引なのかプロンプトなのかが分かります。両方を当てずっぽうで触らずに済みます。
- 根拠のない答えを見つける — 忠実度の指標が、回答に含まれる各主張を実際に引いてきた文書と突き合わせます。読み心地は良いのに根拠がない、という失敗がこれで見えます。
- 生成した試験データから始める — 手持ちの文書から質問を作れるため、誰かが試験項目を手書きする前に最初の評価を実施できます。
- 数値にならないものを判定する — この返答は失礼ではないか、医療上の助言をしていないか、といった文章で書いた基準をモデルが合否判定し、時系列で追えます。
- 回答だけでなくツールの使い方も採点する — エージェント向けの指標では、適切なツールを呼べたか、依頼された内容を達成できたかを見ます。最後の文章だけを見るわけではありません。
Ragasを選ぶ前に
- 多くの指標はモデルの呼び出しです。評価の実行には費用がかかり、同じ入力を2回採点しても数値は一致しません。小さな改善を比較する場面では効いてきます。
- リポジトリは2026年初頭以降ほとんど更新されておらず、所有者も移っています。リリース手順の評価関門として採用する前に、保守の状況を確認してください。
よくある質問
Ragasは商用利用できますか?
RagasはApache-2.0ライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
Ragasはどの形で使えますか?
Ragasはローカル実行・セルフホストの形で利用できます。
ドキュメント
vibrantlabsai/ragas のREADMEより転載(Apache-2.0)。 原文を読む ↗
Objective metrics, intelligent test generation, and data-driven insights for LLM apps
Ragas is your ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. Say goodbye to time-consuming, subjective assessments and hello to data-driven, efficient evaluation workflows. Don’t have a test dataset ready? We also do production-aligned test set generation.
Key Features
- 🎯 Objective Metrics: Evaluate your LLM applications with precision using both LLM-based and traditional metrics.
- 🧪 Test Data Generation: Automatically create comprehensive test datasets covering a wide range of scenarios.
- 🔗 Seamless Integrations: Works flawlessly with popular LLM frameworks like LangChain and major observability tools.
- 📊 Build feedback loops: Leverage production data to continually improve your LLM applications.
:shield: Installation
Pypi:
pip install ragas
Alternatively, from source:
pip install git+https://github.com/vibrantlabsai/ragas
:fire: Quickstart
Clone a Complete Example Project
The fastest way to get started is to use the ragas quickstart command:
# List available templates
ragas quickstart
# Create a RAG evaluation project
ragas quickstart rag_eval
# Specify where you want to create it.
ragas quickstart rag_eval -o ./my-project
Available templates:
rag_eval- Evaluate RAG systems
Coming Soon:
agent_evals- Evaluate AI agentsbenchmark_llm- Benchmark and compare LLMsprompt_evals- Evaluate prompt variationsworkflow_eval- Evaluate complex workflows
Evaluate your LLM App
ragas comes with pre-built metrics for common evaluation tasks. For example, Aspect Critique evaluates any aspect of your output using DiscreteMetric:
import asyncio
from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factory
# Setup your LLM
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)
# Create a custom aspect evaluator
metric = DiscreteMetric(
name="summary_accuracy",
allowed_values=["accurate", "inaccurate"],
prompt="""Evaluate if the summary is accurate and captures key information.
Response: {response}
Answer with only 'accurate' or 'inaccurate'."""
)
# Score your application's output
async def main():
score = await metric.ascore(
llm=llm,
response="The summary of the text is..."
)
print(f"Score: {score.value}") # 'accurate' or 'inaccurate'
print(f"Reason: {score.reason}")
if __name__ == "__main__":
asyncio.run(main())
Note: Make sure your
OPENAI_API_KEYenvironment variable is set.
Find the complete Quickstart Guide
Want help in improving your AI application using evals?
In the past 2 years, we have seen and helped improve many AI applications using evals. If you want help with improving and scaling up your AI application using evals.
🔗 Book a slot or drop us a line: founders@vibrantlabs.com.
🫂 Community
If you want to get more involved with Ragas, check out our discord server. It’s a fun community where we geek out about LLM, Retrieval, Production issues, and more.
🔍 Open Analytics
At Ragas, we believe in transparency. We collect minimal, anonymized usage data to improve our product and guide our development efforts.
✅ No personal or company-identifying information
✅ Open-source data collection code
✅ Publicly available aggregated data
To opt-out, set the RAGAS_DO_NOT_TRACK environment variable to true.
Cite Us
@misc{ragas2024,
author = {VibrantLabs},
title = {Ragas: Supercharge Your LLM Application Evaluations},
year = {2024},
howpublished = {\url{https://github.com/vibrantlabsai/ragas}},
}