Chroma vs Qdrant

Choose a vector database for semantic retrieval in an agent or RAG application.

At a glance

At a glanceChromaQdrant
LicenseApache-2.0Apache-2.0
LanguagesRust, PythonRust
DeploymentRuns locally / Self-hosted / Managed cloudSelf-hosted / Runs locally / Managed cloud
MaturityEstablishedEstablished
Stars29.2k34.2k
Star growth over the last 7 days+58 ★+133 ★
Forks2.5k2.6k
Open issues810717
Last commit28 Aug 202628 Aug 2026
ActivityActiveActive

What each one does

Chroma

The lowest-friction way to get retrieval working: no server to stand up, no cluster to size. That makes it excellent for prototypes and small production loads, and the point at which you outgrow it is a real consideration to plan for rather than discover.

Full entry →

Qdrant

Written in Rust and built so that metadata filters apply during search rather than after it — the difference that keeps results correct when an agent queries a narrow slice of a large collection. Runs from a single container for local development up to a distributed cluster.

Full entry →

What you can do

Chroma

  • Get retrieval running in one file — chromadb.Client() creates an in-memory store and collection.add() handles tokenization, embedding and indexing, so no embedding model is wired up by hand.
  • Filter alongside the vector search — collection.query() takes where for metadata equality and where_document with $contains for substring matching in the same call as the similarity search.
  • Move to client-server mode — chroma run --path /chroma_db_path starts a server against the same four-function API once a single embedded process is no longer enough.
  • Query from Python or JavaScript — The Python client comes from pip install chromadb and the JavaScript one from npm install chromadb; pointed at the same chroma run server, both work on one set of collections, so the job that writes and the service that queries need not share a language.

Qdrant

  • Filter without losing recall — Payload conditions are applied inside the vector search itself rather than as a post-filter over the returned results, and a filter combines keyword, full-text, numeric-range and geo conditions under must, should and must_not.
  • Start from a single container — docker run -p 6333:6333 qdrant/qdrant is the whole local setup, and the README flags that this default has no authentication and is open on every interface; production scales out with sharding and replication, resizing collections with no downtime.
  • Combine dense and sparse in one query — Dense vectors, sparse vectors and multivectors for late-interaction models such as ColBERT can be queried together, with the result lists merged by Reciprocal Rank Fusion or Distribution-Based Score Fusion.
  • Connect from six official clients — Official clients exist for Python (qdrant-client), JavaScript/TypeScript (@qdrant/js-client-rest), Go, Rust, .NET/C# and Java, over a REST API with an OpenAPI 3.0 spec plus a gRPC interface for production-tier traffic.
  • Cut RAM before adding nodes — Built-in quantization lets you trade search speed and precision against memory, with on-disk storage for the vectors that need not stay resident.