Local Vector Search with ChromaDB: Embeddings, Queries, LangChain and a RAG Chatbot

Key takeaways

ChromaDB is an embeddable vector database you can run in-process with pip. This guide covers what it stores, how queries and distances behave, the persistence and embedding-model traps that break retrieval quality, and a small LangChain RAG chain on top.

What ChromaDB is, and what it is not

ChromaDB is an open-source vector database designed to be embedded in a Python (or JavaScript) application. You pip install chromadb, create a client, and you have a working similarity search over text with no separate server, no account and no API key. Under the hood each collection keeps an approximate nearest neighbor index (HNSW) over the vectors, plus a store for the original documents and their metadata, so a query returns text you can show to a user or pass to an LLM, not just ids.

That makes it an easy starting point for retrieval-augmented generation (RAG), semantic search in an internal tool, or de-duplicating similar support tickets. It is also a common default in LangChain and LlamaIndex tutorials. What it is not is a replacement for a replicated, multi-tenant database cluster. Chroma can run as a client/server process, but a single local Chroma instance is still one process with one data directory. Plan backups and capacity the way you would for a SQLite file, not the way you would for a managed service.

Everything below was run against chromadb 1.5.9 on Python 3.11. The API changed noticeably between 0.3, 0.4 and 1.x, so check your installed version (chromadb.__version__) before copying code from older tutorials.

Installation and a first collection

pip install chromadb
import chromadb

client = chromadb.Client()               # in-memory, gone when the process exits
collection = client.create_collection(name="docs")

collection.add(
    documents=[
        "Python is a programming language",
        "Bananas are yellow fruit",
        "pip installs Python packages",
    ],
    metadatas=[
        {"source": "a", "year": 2020},
        {"source": "b", "year": 2021},
        {"source": "a", "year": 2023},
    ],
    ids=["1", "2", "3"],
)

results = collection.query(query_texts=["how do I install packages"], n_results=2)
print(results["ids"])        # [['3', '1']]
print(results["distances"])  # [[1.11..., 1.74...]]

Several details here are worth understanding before building on them:

  • You passed text, not vectors. When a collection has no embedding function configured, Chroma uses its default: the all-MiniLM-L6-v2 sentence-transformer model run through ONNX Runtime, which produces 384-dimensional vectors. The first call downloads the model (about 79 MB) into ~/.cache/chroma/onnx_models/. On a machine without internet access, that first add or query fails, which surprises people in CI and in locked-down containers.
  • Results are nested lists. query accepts several query texts at once, so results["ids"][0] is the list of hits for the first query. Forgetting the [0] is the most common bug in Chroma code.
  • Lower distance means more similar. The values are distances, not scores. With the default metric they are not bounded to 0..1.
  • n_results is a maximum. Asking for 10 results from a 3-document collection returns 3, not an error.
  • Names are validated. Collection names must be 3 to 512 characters from [a-zA-Z0-9._-] and start and end with a letter or digit. create_collection("x") raises InvalidArgumentError.

Distance metrics and why they matter

A new collection uses squared L2 (Euclidean) distance by default. You can see this in the collection’s configuration:

print(collection.configuration_json["hnsw"]["space"])   # 'l2'

Many embedding models, including most OpenAI and sentence-transformer models, are trained so that cosine similarity is the meaningful comparison. For vectors normalized to unit length, L2 and cosine produce the same ranking, so the default often works. For unnormalized vectors they can disagree. Set the metric when you create the collection, because it is fixed once the index exists:

cos = client.create_collection(
    name="docs_cosine",
    metadata={"hnsw:space": "cosine"},   # also: "l2", "ip" (inner product)
)
cos.add(documents=["Python is a programming language"], ids=["1"])
print(cos.query(query_texts=["Python language"], n_results=1)["distances"])
# [[0.12...]]  -> cosine distance = 1 - cosine similarity

With cosine, a distance near 0 means nearly identical direction and 1 means unrelated. That scale is easier to reason about when you want a cut-off such as “ignore anything with distance above 0.5”. Pick the threshold by looking at real queries, though. It depends on the model and the domain, and a value copied from someone else’s project is rarely right.

Using a different embedding model

The default model is small and English-focused. For multilingual text or better retrieval quality, you can plug in another embedding function:

from chromadb.utils import embedding_functions

openai_ef = embedding_functions.OpenAIEmbeddingFunction(
    api_key="sk-...",                   # better: read from an environment variable
    model_name="text-embedding-3-small",
)

collection = client.create_collection(name="docs_openai", embedding_function=openai_ef)

In Chroma 1.x the embedding function is recorded in the collection’s configuration. If you later open the collection with a different one, Chroma refuses instead of silently mixing vector spaces:

ValueError: An embedding function already exists in the collection configuration, and a new one is provided.
If this is intentional, please embed documents separately.
Embedding function conflict: new: my-ef vs persisted: default

That error is doing you a favor. In older versions, nothing stopped you from indexing with one model and querying with another. The result was a collection that returned confident-looking but meaningless matches. Dimensions are checked too: adding a 2-dimensional vector to a collection built on a 384-dimensional model raises InvalidArgumentError: Collection expecting embedding with dimension of 384, got 2.

When you change embedding models, create a new collection and re-embed everything. There is no in-place migration, because vectors from two models do not share a coordinate system.

Filtering with metadata and document content

A vector search finds text that sounds similar, but it does not know that a document is outdated or belongs to another customer. Metadata filters cover that part. where filters on metadata fields; where_document filters on the stored text itself:

collection.query(
    query_texts=["python"],
    n_results=5,
    where={"$and": [{"source": "a"}, {"year": {"$gte": 2021}}]},
)["ids"]
# [['3']]

collection.query(
    query_texts=["python"],
    n_results=5,
    where_document={"$contains": "pip"},
)["ids"]
# [['3']]

Supported operators include $eq, $ne, $gt, $gte, $lt, $lte, $in and $nin, combined with $and and $or. Metadata values must be scalars (strings, numbers, booleans), so store a list of tags as a delimited string or as several boolean fields.

A failure mode I have seen repeatedly in RAG prototypes: someone adds a strict filter such as a tenant id or a date range, the filter matches nothing because of a typo or a type mismatch ("2023" stored as a string, compared with 2023 as an int), and the query returns an empty list. Nothing raises an error. The LLM then gets an empty context and answers from its own training data, which looks plausible. Log how many chunks retrieval returned for every request, and treat zero as a warning. collection.get(where=...) with the same filter is a quick way to check whether the filter matches anything at all.

add, upsert, update and delete

In Chroma 1.5.9, add with an id that already exists does not raise and does not overwrite. The original record stays:

collection.add(documents=["dup"], ids=["1"])
print(collection.get(ids=["1"])["documents"])    # ['Python is a programming language']

collection.upsert(documents=["Python is great"], ids=["1"])
print(collection.get(ids=["1"])["documents"])    # ['Python is great']

This is a real trap for re-ingestion scripts. If you edit a source document and run your ingest job again with add, the collection quietly keeps serving the old text. Use upsert for ingestion jobs, derive ids deterministically from the source (file path plus chunk index, or a content hash), and use delete(where={"source": ...}) to remove chunks for a document that shrank or was deleted. Otherwise stale chunks with old ids stay in the collection.

Persistence

chromadb.Client() keeps everything in memory. For anything you want to keep, use a persistent client:

client = chromadb.PersistentClient(path="./chroma_db")
notes = client.get_or_create_collection("notes")
notes.add(documents=["hello world"], ids=["h"])

# later, in another process
client = chromadb.PersistentClient(path="./chroma_db")
notes = client.get_collection("notes")
print(notes.count())   # 1

Two practical notes. First, treat the directory as owned by one process at a time. If several workers need the same data, run the Chroma server (chroma run --path ./chroma_db) and connect with chromadb.HttpClient(host=..., port=...) instead of pointing many PersistentClient instances at one folder. Second, back up the directory with the application stopped, or snapshot the volume, and keep a note of which embedding model and chunking settings produced it. Those settings are part of the data, and you cannot reproduce the index without them.

LangChain integration

The LangChain wrapper lives in the separate langchain-chroma package. The old from langchain.vectorstores import Chroma import is deprecated.

pip install langchain-chroma langchain-text-splitters langchain-community langchain-openai
from langchain_chroma import Chroma
from langchain_community.document_loaders import TextLoader
from langchain_openai import OpenAIEmbeddings
from langchain_text_splitters import RecursiveCharacterTextSplitter

documents = TextLoader("guide.txt", encoding="utf-8").load()

splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)
chunks = splitter.split_documents(documents)   # metadata such as "source" is kept

vectorstore = Chroma.from_documents(
    documents=chunks,
    embedding=OpenAIEmbeddings(model="text-embedding-3-small"),
    collection_name="guide",
    persist_directory="./chroma_db",
)

for doc, distance in vectorstore.similarity_search_with_score("installing packages", k=3):
    print(round(distance, 3), doc.metadata["source"], doc.page_content[:80])

Chunking deserves more attention than it usually gets. Chunks that are too large put several topics into one vector, so the embedding describes none of them well. Chunks that are too small lose the context that made a sentence meaningful, such as which product or function it was about. The overlap keeps a sentence that crosses a chunk boundary findable from both sides. Start with a few hundred to a thousand characters, then look at what retrieval actually returns for real questions before tuning.

A minimal RAG chain

RetrievalQA from older LangChain examples is deprecated. The current style composes the retriever, a prompt and a model with LCEL:

from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_openai import ChatOpenAI

retriever = vectorstore.as_retriever(search_kwargs={"k": 4})

prompt = ChatPromptTemplate.from_messages([
    ("system",
     "Answer only from the context below. If the answer is not in the context, "
     "say you don't know.\n\nContext:\n{context}"),
    ("human", "{question}"),
])

def format_docs(docs):
    return "\n\n".join(d.page_content for d in docs)

chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | prompt
    | ChatOpenAI(model="gpt-4o-mini", temperature=0)
    | StrOutputParser()
)

print(chain.invoke("How do I install packages?"))

I checked the structure of this chain with LangChain’s DeterministicFakeEmbedding and FakeListChatModel in place of the OpenAI classes, so it runs without an API key. That is also a sensible way to unit-test your own pipeline.

The instruction to say “I don’t know” matters more than it looks. Without it, the model fills gaps from its own training data, and you lose the main benefit of RAG: answers you can trace back to your documents. Returning the retrieved chunks’ source metadata next to the answer makes that tracing possible for users too.

In my experience, a RAG system that “used to work” and now gives worse answers is rarely the model’s fault. The usual cause is the index: someone re-ran ingestion with different chunk settings into the same collection, so old and new chunks are mixed, or a document was updated but add kept the old version. When quality drops, first look at the actual chunks the retriever returned for a failing question. That is usually faster than changing the prompt.