RAG & Retrieval
Retrieval pipelines that ground an agent in your own documents: chunking, embedding, reranking and citation. Retrieval quality, not model choice, is usually what decides whether answers are trustworthy.
16 projects
RAGFlow is a server you run rather than a library you import: documents go into a dataset, it splits them into chunks, and you can read every chunk on screen and hand-correct the ones that came out wrong before any of it reaches a model, with chat answers citing the chunk they came from. Around that sit an agent canvas and an ingestion pipeline you can rebuild yourself, so retrieval, tools and MCP servers can be wired together without writing code. The optical character recognition, table-structure and layout models that make sense of scans run on PDFs and images only — Word, Excel and PowerPoint files are read structurally, and the docs tell you to convert a DOCX to PDF if you want the visual parser on it. The price of all this is weight: a Docker Compose stack with Elasticsearch, MySQL, MinIO and Redis behind it, 4 cores and 16 GB of RAM as the stated minimum, and parsing that the project's own FAQ concedes is slower than LangChain's. If your documents are already clean text and you want retrieval inside your own Python application, LlamaIndex or Haystack will fit better.
The difference from a text extractor is that structure survives: page layout, reading order, table cells, code blocks and formulas come through as parts of a single document representation rather than as a flattened string, which is what decides whether a table in a report is still a table by the time an agent reads it. Everything can run locally, including in air-gapped environments. The cost of that fidelity is compute — machine-learning models run over each page, so this is a pipeline to budget for, not a fast converter to call inline.
AnythingLLM ships the whole retrieval stack as one installable application: a document collector, an embedding model, a vector database and the chat interface arrive together, so you drop files into a workspace and start asking questions without wiring services to each other. Documents attached to a chat go into the model's context in full by default, and only when they overflow the context window does the app offer to chunk and embed them, after which answers are assembled from a few retrieved snippets. New workspaces run in agent mode, so the same chat window can browse the web, query a SQL database, call MCP servers or run a flow you drew in the visual editor. The trade-off is that this is an application rather than a library — you extend it through its settings, its custom-skill plugin format and its REST API, never by importing it into a program of your own — and the three ways to run it are not the same product: the desktop build has no accounts, the Docker build has no bundled model, and the hosted tier has neither MCP nor custom skills.
Where most frameworks start from the agent loop, this one starts from the data: ingestion, indexing strategies, and the retrieval patterns that decide whether answers are actually grounded. Reach for it when the hard part of your problem is the corpus rather than the reasoning.
One deployment puts OpenAI, Anthropic, Google, Bedrock and any OpenAI-compatible endpoint behind a single login, and its agent builder lets non-programmers assemble assistants — instructions, tools, file search, code execution, MCP servers — then share them with access controls by user, group or role, with LDAP, OIDC and SAML covering enterprise sign-on. The cost is operational: a default install already runs six containers, and code execution and web search each add further services or external keys. AnythingLLM remains the simpler pick for one person wanting one container; LibreChat is built for the multi-user case. ClickHouse acquired the project in November 2025, stating the licence stays MIT.
At indexing time an LLM reads every chunk and extracts entities and relations into a knowledge graph kept alongside vector embeddings; at query time five modes choose the path — local for facts about one entity, global for themes spanning documents, hybrid and mix to combine them, naive for plain chunk retrieval. A bundled server adds a REST API, a dashboard with graph visualisation, and an Ollama-compatible endpoint that chat frontends can talk to as if it were a model. The cost structure is the decision: building the graph spends an LLM call on every chunk, the project names a 30B-class model as the local minimum for extraction, and the out-of-the-box storage is stated to be unsuitable for production — if cheap chunk search is all you need, the existence of its own naive mode concedes that plainer RAG suffices.
Ordinary retrieval finds the chunks most similar to a question, which fails when the answer is not written down in any one chunk — what the main themes are, how two people are connected. GraphRAG instead extracts entities, relationships and claims from every passage, clusters the resulting graph into communities and summarises each one, then answers from those summaries. The price is paid at indexing time: a model runs over every text unit and again over every community, so building the index scales with the size of the corpus, not with how often you query it.
Vector search finds passages that resemble the question, and resemblance is not relevance — on a long regulatory filing the paragraph that answers a question often shares almost no vocabulary with it. PageIndex drops the vector store. It builds a tree of the document's real structure, section by section, and then has a model walk that tree the way a person would use a table of contents: narrowing down, opening the part that should hold the answer, reading it. Nothing is chunked, and every answer points at the section it came from. The trade is per-query cost, since retrieval now involves reasoning rather than a lookup, and the design targets long individual documents rather than a corpus of short ones.
Formerly Danswer, this is the whole surface a company usually assembles by hand — a chat interface, indexing from more than fifty sources, custom agents with their own instructions and actions, a multi-step research mode that returns a report, web search, code execution and artifacts — deployable by one command and pointed at any model provider, self-hosted or proprietary. Two things to weigh: the full deployment is a stack of index, workers, inference servers, cache and blob store, and directories named ee carry a separate enterprise licence rather than MIT.
Most memory layers overwrite: a new fact replaces the old one and last year's answer becomes unrecoverable. Graphiti instead marks the old relationship as no longer holding, with the dates attached, so both what is true now and what was true then remain queryable — and it separates when something happened from when the system learned about it, which is what keeps late-arriving information from rewriting history. Retrieval combines vector similarity, full-text search and graph traversal. The costs are concrete: a graph database has to run alongside your stack, and every ingest calls a model to extract entities and relationships.
Semantic Kernel puts a kernel at the centre of the application: a container you register model connections and plugins into, which the agents you build then draw on. Mark ordinary methods as kernel functions and the model can call them, import a whole API from an OpenAPI description, wrap filters around a call to log, cache, redact or stop it, and hand a conversation between several specialised agents. The trade-off is direction rather than quality. Microsoft has named Agent Framework the successor and says the majority of new features will be built there, so this codebase keeps getting fixes and will see some existing features reach general availability, but few new ideas. It suits a team extending something already written against it far better than a project starting from nothing today.
Haystack builds document-grounded answering and agents out of components — file converters, splitters, embedders, retrievers, chat generators, tools — that you register in a Pipeline and connect output-name to input-name, with a mismatch raising an error at connection time rather than partway through a run. Branches and capped loops live in the same graph, and an Agent is itself a component, so a whole tool-calling loop drops into a pipeline or becomes a tool for another agent. The cost is that the wiring is yours: there is no one-call path from a folder of PDFs to a working assistant, and the library runs inside your own process — serving a pipeline over HTTP or as an MCP server means adding Hayhooks, a separate deepset project, or writing that wrapper yourself.
A vector becomes a column type, and from there everything Postgres already does applies: a nearest-neighbour query is an ORDER BY with a LIMIT, filters are WHERE clauses evaluated by the same planner, the embedding and the row it describes are written in one transaction, and the backups, replicas and roles you already run cover the vectors too. What you give up is the operational apparatus a dedicated engine provides — sharding a collection, isolating tenants — and the fact that index builds now compete with the rest of the workload for the same machine.
Three things distinguish it from a bare index. Objects are stored with their properties, so a filter is part of the search rather than something applied afterwards; a vectoriser module can generate the embeddings on write, removing the separate pipeline that otherwise drifts out of sync; and multi-tenancy is a first-class construct where each tenant gets its own shard and can be parked or offloaded to object storage. It is a database you run and operate, though — for an embedded index inside one process, this is more machinery than the problem needs.
The failure most retrieval systems have is invisible from the outside — the answer reads well and is not supported by anything that was retrieved. Ragas measures that directly. It splits the pipeline in two and scores each half: whether the retrieved passages actually contained what was needed, and whether the generated answer stayed within them. It can also build a starting test set out of your own documents, so the first evaluation does not wait for someone to hand-write a hundred questions. Nearly every metric is itself a model call, which is the thing to plan around: a full run costs money, and two runs on the same data will not produce identical numbers.
The headline feature is AI Services: you declare a Java interface, annotate it, and the library supplies the implementation that builds the prompt, calls the model and maps the reply back onto your return type. Underneath, a plain ChatModel API is there whenever the declarative layer gets in the way, and RAG, tool calling and integrations for Spring Boot, Quarkus, Micronaut and Helidon are first-party. Despite the name it is not a port of the Python library, so material from that ecosystem does not transfer.