← Back to all projects

GraphRAG

Builds a knowledge graph from your documents so questions about the whole corpus become answerable

OfficialMIT
Stars
35.7k
Forks
3.8k
Open issues
36
Last commit
24 Aug 2026

What is GraphRAG?

Ordinary retrieval finds the chunks most similar to a question, which fails when the answer is not written down in any one chunk — what the main themes are, how two people are connected. GraphRAG instead extracts entities, relationships and claims from every passage, clusters the resulting graph into communities and summarises each one, then answers from those summaries. The price is paid at indexing time: a model runs over every text unit and again over every community, so building the index scales with the size of the corpus, not with how often you query it.

What can you do with GraphRAG?

  • Answer questions no single passage contains — Global search works from the community summaries rather than from retrieved chunks, which is what makes corpus-wide questions — the recurring themes, what changed across a set of reports — answerable at all.
  • Start from an entity and walk outwards — Local search anchors on a specific entity and explores its neighbours and associated concepts, so a question about one person, product or place gathers the scattered mentions instead of the most similar paragraph.
  • Inspect the graph it derived — Indexing produces entities, relationships and claims as data you can look at, so a wrong answer can be traced to a wrong extraction rather than disappearing into an embedding.
  • Summaries at every level of the hierarchy — The Leiden clustering method groups the graph into communities and the summaries are generated bottom-up, giving both a fine-grained and a whole-corpus view from the same index.
  • Tune the extraction to your domain — The prompts that decide what counts as an entity or a claim are meant to be adapted; the documentation treats prompt tuning as part of setup rather than an optimisation to do later.

Before you choose GraphRAG

  • A model runs over every text unit during extraction and again over every community during summarisation, so index-building cost tracks the size of the corpus rather than the number of questions you will ask it.
  • The documentation frames it as the answer where baseline retrieval struggles to connect the dots; for straightforward lookups in a modest corpus, ordinary retrieval remains the cheaper and simpler choice.

Star history

21 Aug to 28 Aug · +106

35.6k35.7k

Frequently asked questions

Is GraphRAG free for commercial use?

GraphRAG is released under the MIT licence — OSI-approved open source, which permits commercial use.

How can GraphRAG be deployed?

GraphRAG is available as Self-hosted / Runs locally.

Documentation

Reproduced from the microsoft/graphrag README, published under MIT. Read the original ↗

GraphRAG

[!WARNING] GraphRAG is a research project that explores the functional use of graphs to form a targeted context for question answering. Since our first release in July 2024 the capabilities of frontier models have changed dramatically, and our portfolio of research projects has diversified to match. This project is largely in maintenance mode, and won’t be accepting new PRs or implementing new features. We’ll perform bug fixes and dependency updates as appropriate, particularly to address CVEs as they arise.

👉 Microsoft Research Blog Post 👉 Read the docs 👉 GraphRAG Arxiv

Overview

The GraphRAG project is a data pipeline and transformation suite that is designed to extract meaningful, structured data from unstructured text using the power of LLMs.

To learn more about GraphRAG and how it can be used to enhance your LLM’s ability to reason about your private data, please visit the Microsoft Research Blog Post.

Quickstart

To get started with the GraphRAG system we recommend trying the command line quickstart.

Repository Guidance

This repository presents a methodology for using knowledge graph memory structures to enhance LLM outputs. Please note that the provided code serves as a demonstration and is not an officially supported Microsoft offering.

⚠️ Warning: GraphRAG indexing can be an expensive operation, please read all of the documentation to understand the process and costs involved, and start small.

Diving Deeper

Prompt Tuning

Using GraphRAG with your data out of the box may not yield the best possible results. We strongly recommend to fine-tune your prompts following the Prompt Tuning Guide in our documentation.

Versioning

Please see the breaking changes document for notes on our approach to versioning the project.

Always run graphrag init --root [path] --force between minor version bumps to ensure you have the latest config format. Run the provided migration notebook between major version bumps if you want to avoid re-indexing prior datasets. Note that this will overwrite your configuration and prompts, so back them up if necessary.

Responsible AI FAQ

See RAI_TRANSPARENCY.md

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft’s Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party’s policies.

Privacy

Microsoft Privacy Statement