← プロジェクト一覧に戻る

Ragas

検索を使う仕組みを、必要な文書を引けたかと、答えがその文書の中に収まっているかの2点で採点する

Apache-2.0
スター
15.5k
フォーク
1.7k
オープンIssue
577
最終コミット
2026年2月24日

Ragasとは

検索を使う仕組みで最も多い失敗は、外から見えません。答えは読みやすいのに、引いてきた文書のどこにも根拠がない、という状態です。Ragasはそれを直接測ります。処理を2つに分け、必要な内容を含む文書を引けていたか、生成された答えがその範囲に収まっているかを別々に採点します。手持ちの文書から試験用のデータを作ることもできるため、最初の評価のために質問を100件手書きするところから始めずに済みます。ほとんどの指標自体がモデルの呼び出しである点は、あらかじめ織り込んでおくべきです。一通り走らせれば費用がかかり、同じデータで2回実行しても数値は完全には一致しません。

Ragasで何ができますか?

  • 検索の失敗と生成の失敗を切り分ける — 検索と回答を別々に採点するため、直すべきなのが索引なのかプロンプトなのかが分かります。両方を当てずっぽうで触らずに済みます。
  • 根拠のない答えを見つける — 忠実度の指標が、回答に含まれる各主張を実際に引いてきた文書と突き合わせます。読み心地は良いのに根拠がない、という失敗がこれで見えます。
  • 生成した試験データから始める — 手持ちの文書から質問を作れるため、誰かが試験項目を手書きする前に最初の評価を実施できます。
  • 数値にならないものを判定する — この返答は失礼ではないか、医療上の助言をしていないか、といった文章で書いた基準をモデルが合否判定し、時系列で追えます。
  • 回答だけでなくツールの使い方も採点する — エージェント向けの指標では、適切なツールを呼べたか、依頼された内容を達成できたかを見ます。最後の文章だけを見るわけではありません。

Ragasを選ぶ前に

  • 多くの指標はモデルの呼び出しです。評価の実行には費用がかかり、同じ入力を2回採点しても数値は一致しません。小さな改善を比較する場面では効いてきます。
  • リポジトリは2026年初頭以降ほとんど更新されておらず、所有者も移っています。リリース手順の評価関門として採用する前に、保守の状況を確認してください。

よくある質問

Ragasは商用利用できますか?

RagasはApache-2.0ライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

Ragasはどの形で使えますか?

Ragasはローカル実行・セルフホストの形で利用できます。

ドキュメント

vibrantlabsai/ragas のREADMEより転載(Apache-2.0)。 原文を読む ↗

Objective metrics, intelligent test generation, and data-driven insights for LLM apps

Ragas is your ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. Say goodbye to time-consuming, subjective assessments and hello to data-driven, efficient evaluation workflows. Don’t have a test dataset ready? We also do production-aligned test set generation.

Key Features

  • 🎯 Objective Metrics: Evaluate your LLM applications with precision using both LLM-based and traditional metrics.
  • 🧪 Test Data Generation: Automatically create comprehensive test datasets covering a wide range of scenarios.
  • 🔗 Seamless Integrations: Works flawlessly with popular LLM frameworks like LangChain and major observability tools.
  • 📊 Build feedback loops: Leverage production data to continually improve your LLM applications.

:shield: Installation

Pypi:

pip install ragas

Alternatively, from source:

pip install git+https://github.com/vibrantlabsai/ragas

:fire: Quickstart

Clone a Complete Example Project

The fastest way to get started is to use the ragas quickstart command:

# List available templates
ragas quickstart

# Create a RAG evaluation project
ragas quickstart rag_eval

# Specify where you want to create it.
ragas quickstart rag_eval -o ./my-project

Available templates:

  • rag_eval - Evaluate RAG systems

Coming Soon:

  • agent_evals - Evaluate AI agents
  • benchmark_llm - Benchmark and compare LLMs
  • prompt_evals - Evaluate prompt variations
  • workflow_eval - Evaluate complex workflows

Evaluate your LLM App

ragas comes with pre-built metrics for common evaluation tasks. For example, Aspect Critique evaluates any aspect of your output using DiscreteMetric:

import asyncio
from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factory

# Setup your LLM
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)

# Create a custom aspect evaluator
metric = DiscreteMetric(
    name="summary_accuracy",
    allowed_values=["accurate", "inaccurate"],
    prompt="""Evaluate if the summary is accurate and captures key information.

Response: {response}

Answer with only 'accurate' or 'inaccurate'."""
)

# Score your application's output
async def main():
    score = await metric.ascore(
        llm=llm,
        response="The summary of the text is..."
    )
    print(f"Score: {score.value}")  # 'accurate' or 'inaccurate'
    print(f"Reason: {score.reason}")


if __name__ == "__main__":
    asyncio.run(main())

Note: Make sure your OPENAI_API_KEY environment variable is set.

Find the complete Quickstart Guide

Want help in improving your AI application using evals?

In the past 2 years, we have seen and helped improve many AI applications using evals. If you want help with improving and scaling up your AI application using evals.

🔗 Book a slot or drop us a line: founders@vibrantlabs.com.

🫂 Community

If you want to get more involved with Ragas, check out our discord server. It’s a fun community where we geek out about LLM, Retrieval, Production issues, and more.

🔍 Open Analytics

At Ragas, we believe in transparency. We collect minimal, anonymized usage data to improve our product and guide our development efforts.

✅ No personal or company-identifying information

✅ Open-source data collection code

✅ Publicly available aggregated data

To opt-out, set the RAGAS_DO_NOT_TRACK environment variable to true.

Cite Us

@misc{ragas2024,
  author       = {VibrantLabs},
  title        = {Ragas: Supercharge Your LLM Application Evaluations},
  year         = {2024},
  howpublished = {\url{https://github.com/vibrantlabsai/ragas}},
}