Vector & Storage

Storage engines for embeddings. Beyond raw search speed, look at metadata filtering, hybrid search and what happens to recall when the collection grows past what fits in memory.

7 projects

ClickHouseObservability

A column-oriented engine for append-heavy, high-cardinality data, which is the shape agent telemetry takes: one row per model call, filtered later by model, cost or error. Langfuse moved its traces here from Postgres in December 2024 and ClickHouse acquired the project in January 2026, so a self-hosted observability stack increasingly means running this underneath. It is the wrong choice for anything transactional — updates and deletes rewrite whole column parts rather than edit rows — so treat it as where agent runs land, not as the database your application writes state to.

OfficialApache-2.049.5k+133
MilvusVector & Storage

Separates storage from compute so indexing and querying scale independently — the architecture you want at billions of vectors. That capability comes with operational weight; below a few million vectors, something simpler will serve you better.

OfficialApache-2.045.9k+124
FaissVector & Storage

It is built around an index that holds a set of vectors and searches them, and the dozens of available index structures exist because the trade-offs are real: search time against result quality against memory per vector against how long the index takes to build, and whether it needs training data at all. Some methods keep only a compressed representation and never the original vectors, which is how a single server reaches billions of them. There is no server, no filtering and no access control here — those are what the databases built on top of it add.

OfficialMIT40.8k+26
QdrantVector & Storage

Written in Rust and built so that metadata filters apply during search rather than after it — the difference that keeps results correct when an agent queries a narrow slice of a large collection. Runs from a single container for local development up to a distributed cluster.

OfficialApache-2.034.2k+133
ChromaVector & Storage

The lowest-friction way to get retrieval working: no server to stand up, no cluster to size. That makes it excellent for prototypes and small production loads, and the point at which you outgrow it is a real consideration to plan for rather than discover.

OfficialApache-2.029.2k+58
pgvectorVector & Storage

A vector becomes a column type, and from there everything Postgres already does applies: a nearest-neighbour query is an ORDER BY with a LIMIT, filters are WHERE clauses evaluated by the same planner, the embedding and the row it describes are written in one transaction, and the backups, replicas and roles you already run cover the vectors too. What you give up is the operational apparatus a dedicated engine provides — sharding a collection, isolating tenants — and the fact that index builds now compete with the rest of the workload for the same machine.

OfficialPostgreSQL License22.8k+108
WeaviateVector & Storage

Three things distinguish it from a bare index. Objects are stored with their properties, so a filter is part of the search rather than something applied afterwards; a vectoriser module can generate the embeddings on write, removing the separate pipeline that otherwise drifts out of sync; and multi-tenancy is a first-class construct where each tenant gets its own shard and can be parked or offloaded to object storage. It is a database you run and operate, though — for an embedded index inside one process, this is more machinery than the problem needs.

OfficialBSD-3-Clause16.8k+14