AI/LLM

Build a RAG pipeline from scratch — the actually simple version

Build a working RAG pipeline in one Python file — no LangChain, no vector DB signup. The 4 steps that actually matter, with real runnable code.

Every RAG tutorial online starts with LangChain, LlamaIndex, or a hosted vector database signup. None of that is necessary to understand what a RAG pipeline actually is. This walkthrough builds one in a single Python file with three libraries — no framework, no cloud accounts, no config files.

What RAG actually is

RAG (Retrieval-Augmented Generation) is a pattern where an LLM answers a question using text pulled from a private document store at query time, not from what the model memorized during training. It exists because pretrained LLMs don’t know a team’s docs, a product’s changelog, or last week’s meeting notes, and fine-tuning every time the source changes is impractical.

When RAG is worth adding

RAG makes sense when at least one of these is true:

  • Source data changes often (docs, tickets, chat logs, code)
  • Source data is private (company internals, customer records)
  • Source data is large — thousands of pages — and answers must cite specific passages

If the entire dataset fits in a single 200K-token context window (roughly 500 pages) and doesn’t change between requests, RAG is often overkill — sending the whole corpus with the question is simpler and sometimes cheaper. That trade-off is worth its own walkthrough, coming later in this series.

The 4-step pipeline

Every RAG system, regardless of framework or vendor, does some version of these four steps:

  1. Chunk — split source documents into smaller passages
  2. Embed — convert each chunk into a vector (a list of floats)
  3. Store — save chunks + vectors in a vector database
  4. Retrieve + generate — embed the user’s question, find the closest chunks, hand them to an LLM

Everything else — reranking, hybrid search, query rewriting, evaluation — is optional refinement on top of this loop.

Setup

Three libraries, one pip install:

pip install chromadb sentence-transformers anthropic
  • chromadb — a small local vector database (no server, no signup, in-memory or disk-backed)
  • sentence-transformers — the embedding model, downloaded once on first run
  • anthropic — client for Claude, used for the answer step

The embedding model (all-MiniLM-L6-v2) has 22 million parameters (~90 MB download) and runs on CPU in milliseconds. The Claude client needs an API key from console.anthropic.com — Haiku 4.5 is $1 per million input tokens, so a full test session costs cents.

Set the key as an env var:

export ANTHROPIC_API_KEY=sk-ant-...

Step 1 — Chunk the documents

A “chunk” is a small passage the embedding model can turn into one vector. Chunks that are too big (thousands of tokens) blur meaning across topics. Too small (single sentences) lose surrounding context. A good starting point: 200-400 words per chunk with a 40-word overlap between adjacent chunks.

For this tutorial the source is a short list of imaginary product-doc passages held in a Python list:

docs = [
    "Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.",
    "Aurora API authentication uses a bearer token in the Authorization header.",
    "Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.",
    "Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.",
    "Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.",
    "Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.",
]

Each entry is already a self-contained chunk, so no splitting is needed here. For real documents (Markdown files, PDFs, HTML pages), a character-based splitter works for a first pass:

def chunk_text(text: str, size: int = 1500, overlap: int = 200) -> list[str]:
    """Split text into overlapping character windows. Cheap and predictable."""
    chunks = []
    start = 0
    while start < len(text):
        chunks.append(text[start:start + size])
        start += size - overlap
    return chunks

Character splitting is imperfect — it can cut mid-sentence — but for a first pipeline it beats spending an hour picking the “right” splitter. Better strategies (heading-aware, semantic, sentence-boundary) get their own deep-dive later in this series.

Step 2 — Embed the chunks

An embedding is a fixed-length vector that represents the meaning of a text passage. Chunks with similar meanings sit close together in vector space; unrelated chunks are far apart. That distance is the entire mechanism behind retrieval.

from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedder.encode(docs).tolist()

all-MiniLM-L6-v2 outputs 384-dimensional vectors. The .tolist() call converts the NumPy array into plain Python lists so Chroma can serialize them.

First run downloads the model to ~/.cache/huggingface/ — a one-time 22 MB download.

Step 3 — Store chunks and embeddings in a vector database

Chroma runs in-process. EphemeralClient() keeps everything in RAM (fine for tutorials and tests). PersistentClient(path="./chroma_data") writes to disk and survives restarts.

import chromadb

client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")

collection.add(
    ids=[f"doc_{i}" for i in range(len(docs))],
    documents=docs,
    embeddings=embeddings,
)

Every stored item needs a unique id. Chroma keeps the raw document text alongside the embedding, so retrieval later returns both the vector match and the original text to feed to the LLM.

Step 4 — Retrieve, then generate the answer

Retrieval is symmetric with storage: embed the user’s question with the same model, ask Chroma for the closest N chunks, and hand them to the LLM as context.

import anthropic

def answer(question: str, k: int = 3) -> str:
    q_embedding = embedder.encode([question]).tolist()

    results = collection.query(
        query_embeddings=q_embedding,
        n_results=k,
    )
    context = "\n\n".join(results["documents"][0])

    llm = anthropic.Anthropic()
    reply = llm.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=400,
        messages=[{
            "role": "user",
            "content": (
                "Answer the question using ONLY the context below. "
                "If the answer is not in the context, say so.\n\n"
                f"Context:\n{context}\n\n"
                f"Question: {question}"
            ),
        }],
    )
    return reply.content[0].text

Two details worth pinning down:

  • Use the same embedding model for storage and retrieval. Mixing models silently breaks similarity math — vectors from two different models are not directly comparable, and search results become noise.
  • Instruct the model to say “I don’t know.” Without that clause, LLMs will invent an answer from prior training when retrieval misses. It is one of the cheapest reliability wins in RAG.

The full script

Everything above, in one file (rag.py):

import chromadb
import anthropic
from sentence_transformers import SentenceTransformer

docs = [
    "Aurora API rate limits: free tier is 60 requests per minute. Paid tiers start at 600 rpm.",
    "Aurora API authentication uses a bearer token in the Authorization header.",
    "Aurora supports Python, Go, and Node.js SDKs. Community SDKs exist for Ruby and Rust.",
    "Aurora billing is per-request. Failed requests (5xx) are not charged. 4xx requests are charged.",
    "Aurora response latency: p50 is 120ms, p99 is 800ms across all regions.",
    "Aurora data residency: EU customers can pin storage to Frankfurt. US customers use us-east-1 by default.",
]

embedder = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = embedder.encode(docs).tolist()

client = chromadb.EphemeralClient()
collection = client.create_collection(name="aurora_docs")
collection.add(
    ids=[f"doc_{i}" for i in range(len(docs))],
    documents=docs,
    embeddings=embeddings,
)

def answer(question: str, k: int = 3) -> str:
    q_embedding = embedder.encode([question]).tolist()
    results = collection.query(query_embeddings=q_embedding, n_results=k)
    context = "\n\n".join(results["documents"][0])
    llm = anthropic.Anthropic()
    reply = llm.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=400,
        messages=[{
            "role": "user",
            "content": (
                "Answer the question using ONLY the context below. "
                "If the answer is not in the context, say so.\n\n"
                f"Context:\n{context}\n\nQuestion: {question}"
            ),
        }],
    )
    return reply.content[0].text

print(answer("Are failed requests charged?"))

Expected output:

No — failed requests that return a 5xx status are not charged. 4xx requests are charged.

That is a complete RAG loop.

What’s still missing before this is production-grade

The pipeline above works, but a production RAG system adds five things:

  1. A persistent vector store. EphemeralClient() loses data on restart. Swap for PersistentClient(path=...), or graduate to Qdrant, pgvector, or Pinecone when the collection grows past a few hundred thousand chunks. A comparison of vector databases is coming next in this series.
  2. A smarter chunker. Markdown-aware, heading-aware, or semantic splitters preserve context better than fixed-size character windows. A deep-dive on chunking strategies follows.
  3. Metadata filtering. Attaching {"source": "changelog", "date": "2026-08-01"} to each chunk lets retrieval filter by tag or date range, not similarity alone.
  4. Reranking. Pull 20 candidates from the vector DB, then rerank them with a cross-encoder before sending the top 3 to the LLM. Cuts hallucinations and improves precision.
  5. Evaluation. Something that measures whether the pipeline actually returned the right answer, not just “an” answer. Also getting its own article in this series.

Everything past those is optimization. The four-step loop above is the whole idea.

Common mistakes at this stage

  • Storing the whole document as one chunk. Retrieval then returns the entire document every time, and the LLM ignores the middle.
  • Chunking on newlines only. A single paragraph in a well-formatted document can be 2000 words. Fix chunk size in characters or tokens, not lines.
  • Using different embedding models for storage and query. All similarity math becomes noise; results feel random.
  • Sending the retrieved chunks without an instruction to only use them. The LLM blends prior training with the new context and hallucinates confidently.
  • Skipping the “I don’t know” clause. Guarantees a plausible-sounding wrong answer whenever retrieval misses.

Bottom line

RAG is four steps: chunk, embed, store, retrieve+answer. Everything else — LangChain, LlamaIndex, hosted vector DBs, hybrid search, reranking, query rewriting — is scaffolding built on top of this loop. Once the version above works on real documents, layering the extras one at a time reveals exactly what each one changes and why.