← Back to all projects

Ragas

Scores retrieval on two questions: did it find the right passages, and did the answer stay inside them

Apache-2.0
Stars
15.5k
Forks
1.7k
Open issues
577
Last commit
24 Feb 2026

What is Ragas?

The failure most retrieval systems have is invisible from the outside — the answer reads well and is not supported by anything that was retrieved. Ragas measures that directly. It splits the pipeline in two and scores each half: whether the retrieved passages actually contained what was needed, and whether the generated answer stayed within them. It can also build a starting test set out of your own documents, so the first evaluation does not wait for someone to hand-write a hundred questions. Nearly every metric is itself a model call, which is the thing to plan around: a full run costs money, and two runs on the same data will not produce identical numbers.

What can you do with Ragas?

  • Separate a retrieval failure from a generation failure — Retrieval and answer are scored apart, so you learn whether to fix the index or the prompt instead of guessing at both.
  • Catch answers with nothing behind them — The faithfulness score checks each claim in the answer against the passages that were retrieved — the failure that reads perfectly well.
  • Start from a generated test set — Questions can be produced from your own documents, so a first evaluation is possible before anyone has written test cases by hand.
  • Judge what a number cannot — A free-text criterion — is this reply rude, does it give medical advice — is evaluated by a model as a pass or fail you can track over time.
  • Score tool use as well as answers — Metrics for agent behaviour look at whether the right tools were called and whether the run achieved what was asked, not only at the final text.

Before you choose Ragas

  • Most metrics are model calls, so an evaluation run has a bill and a variance — the same inputs scored twice will not give identical numbers, which matters when you are comparing small changes.
  • The repository has been quiet since early 2026 and moved to a new owner, so confirm the maintenance situation before adopting it as the evaluation gate in a release process.

Frequently asked questions

Is Ragas free for commercial use?

Ragas is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.

How can Ragas be deployed?

Ragas is available as Runs locally / Self-hosted.

Documentation

Reproduced from the vibrantlabsai/ragas README, published under Apache-2.0. Read the original ↗

Objective metrics, intelligent test generation, and data-driven insights for LLM apps

Ragas is your ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. Say goodbye to time-consuming, subjective assessments and hello to data-driven, efficient evaluation workflows. Don’t have a test dataset ready? We also do production-aligned test set generation.

Key Features

  • 🎯 Objective Metrics: Evaluate your LLM applications with precision using both LLM-based and traditional metrics.
  • 🧪 Test Data Generation: Automatically create comprehensive test datasets covering a wide range of scenarios.
  • 🔗 Seamless Integrations: Works flawlessly with popular LLM frameworks like LangChain and major observability tools.
  • 📊 Build feedback loops: Leverage production data to continually improve your LLM applications.

:shield: Installation

Pypi:

pip install ragas

Alternatively, from source:

pip install git+https://github.com/vibrantlabsai/ragas

:fire: Quickstart

Clone a Complete Example Project

The fastest way to get started is to use the ragas quickstart command:

# List available templates
ragas quickstart

# Create a RAG evaluation project
ragas quickstart rag_eval

# Specify where you want to create it.
ragas quickstart rag_eval -o ./my-project

Available templates:

  • rag_eval - Evaluate RAG systems

Coming Soon:

  • agent_evals - Evaluate AI agents
  • benchmark_llm - Benchmark and compare LLMs
  • prompt_evals - Evaluate prompt variations
  • workflow_eval - Evaluate complex workflows

Evaluate your LLM App

ragas comes with pre-built metrics for common evaluation tasks. For example, Aspect Critique evaluates any aspect of your output using DiscreteMetric:

import asyncio
from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factory

# Setup your LLM
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)

# Create a custom aspect evaluator
metric = DiscreteMetric(
    name="summary_accuracy",
    allowed_values=["accurate", "inaccurate"],
    prompt="""Evaluate if the summary is accurate and captures key information.

Response: {response}

Answer with only 'accurate' or 'inaccurate'."""
)

# Score your application's output
async def main():
    score = await metric.ascore(
        llm=llm,
        response="The summary of the text is..."
    )
    print(f"Score: {score.value}")  # 'accurate' or 'inaccurate'
    print(f"Reason: {score.reason}")


if __name__ == "__main__":
    asyncio.run(main())

Note: Make sure your OPENAI_API_KEY environment variable is set.

Find the complete Quickstart Guide

Want help in improving your AI application using evals?

In the past 2 years, we have seen and helped improve many AI applications using evals. If you want help with improving and scaling up your AI application using evals.

🔗 Book a slot or drop us a line: founders@vibrantlabs.com.

🫂 Community

If you want to get more involved with Ragas, check out our discord server. It’s a fun community where we geek out about LLM, Retrieval, Production issues, and more.

🔍 Open Analytics

At Ragas, we believe in transparency. We collect minimal, anonymized usage data to improve our product and guide our development efforts.

✅ No personal or company-identifying information

✅ Open-source data collection code

✅ Publicly available aggregated data

To opt-out, set the RAGAS_DO_NOT_TRACK environment variable to true.

Cite Us

@misc{ragas2024,
  author       = {VibrantLabs},
  title        = {Ragas: Supercharge Your LLM Application Evaluations},
  year         = {2024},
  howpublished = {\url{https://github.com/vibrantlabsai/ragas}},
}