
Chroma
Embedded vector store that runs from a single import
Overview
The lowest-friction way to get retrieval working: no server to stand up, no cluster to size. That makes it excellent for prototypes and small production loads, and the point at which you outgrow it is a real consideration to plan for rather than discover.
What can you do with Chroma?
- Get retrieval running in one file —
chromadb.Client()creates an in-memory store andcollection.add()handles tokenization, embedding and indexing, so no embedding model is wired up by hand. - Filter alongside the vector search —
collection.query()takeswherefor metadata equality andwhere_documentwith$containsfor substring matching in the same call as the similarity search. - Move to client-server mode —
chroma run --path /chroma_db_pathstarts a server against the same four-function API once a single embedded process is no longer enough. - Query from Python or JavaScript — The Python client comes from
pip install chromadband the JavaScript one fromnpm install chromadb; pointed at the samechroma runserver, both work on one set of collections, so the job that writes and the service that queries need not share a language.
Documentation
Reproduced from the chroma-core/chroma README, published under Apache-2.0. Read the original ↗

pip install chromadb # python client
# for javascript, npm install chromadb!
# for client-server mode, chroma run --path /chroma_db_path
Chroma Cloud
Our hosted service, Chroma Cloud, powers serverless vector, hybrid, and full-text search. It’s extremely fast, cost-effective, scalable and painless. Create a DB and try it out in under 30 seconds with $5 of free credits.
API
The core API is only 4 functions (run our 💡 Google Colab):
import chromadb
# setup Chroma in-memory, for easy prototyping. Can add persistence easily!
client = chromadb.Client()
# Create collection. get_collection, get_or_create_collection, delete_collection also available!
collection = client.create_collection("all-my-documents")
# Add docs to the collection. Can also update and delete. Row-based API coming soon!
collection.add(
documents=["This is document1", "This is document2"], # we handle tokenization, embedding, and indexing automatically. You can skip that and add your own embeddings as well
metadatas=[{"source": "notion"}, {"source": "google-docs"}], # filter on these!
ids=["doc1", "doc2"], # unique for each doc
)
# Query/search 2 most similar results. You can also .get by id
results = collection.query(
query_texts=["This is a query document"],
n_results=2,
# where={"metadata_field": "is_equal_to_this"}, # optional filter
# where_document={"$contains":"search_string"} # optional filter
)
Learn about all features on our Docs
Get involved
Chroma is a rapidly developing project. We welcome PR contributors and ideas for how to improve the project.
- Join the conversation on Discord -
#contributingchannel - Review the 🛣️ Roadmap and contribute your ideas
- Grab an issue and open a PR -
Good first issue tag - Read our contributing guide
Release Cadence
We currently release new tagged versions of the pypi and npm packages on Mondays. Hotfixes go out at any time during the week.