01 May 2026
Local RAG system on Ollama — interrogating classic literature with ChromaDB + Mistral
RAG sounds complicated until you build it once.
retrieve documents. feed them to the model as context. let the model answer. that's the whole thing.
built a local version — Mistral and nomic-embed-text via Ollama, ChromaDB for the vector store, three books from Project Gutenberg as the corpus. Frankenstein, Dracula, The Time Machine. nothing leaves the machine.
here's what the vector database is actually doing: on first run, each book gets split into ~500 token chunks. nomic-embed-text converts every chunk into a list of numbers — a vector that represents its meaning in high-dimensional space. ChromaDB stores those vectors on disk.
when you ask a question, the same embedding model converts your question into a vector. ChromaDB then finds the chunks whose vectors are closest to it — not keyword matching, semantic similarity. "who created the monster" finds the right passage even if it never uses the word "monster."
those chunks go into Mistral's context window as the answer source. the model reads them, answers from them. it's not recalling from training data — it's reading the book.
subsequent runs skip ingestion entirely. the DB is already on disk. first run is slow. every run after is fast.
classic literature as a corpus was deliberate. public domain, long texts, questions with known answers. if the retrieval fails, the answer fails — which is exactly what you want when you're debugging a pipeline.
same architecture, any text.