Skip to content

Design a RAG pipeline end to end

IntermediateAsked very oftenSystem designDesignRAG
#rag#vector-search#embeddings#reranking#grounding

What interviewers are testing

System-design interviewers use RAG to test whether you can decompose a fuzzy product requirement into two concrete data pipelines and reason about quality at every stage. Strong candidates discuss chunking and embedding-space consistency, explain why retrieval is optimized for recall and reranking for precision, and can name the failure modes — hallucination on retrieval misses, stale indexes, context dilution — with mitigations. Weak answers stop at "embed everything and put it in a vector database".

Mental model

RAG is a retrieval contract between two pipelines: ingestion turns documents into retrievable embedded chunks with metadata, and query turns a question into an embedded search, a reranked shortlist, and a grounded, cited answer. Retrieval quality sets the ceiling for generation quality — if the right passage never makes the shortlist, no prompt can recover the answer, so evaluate retrieval and generation separately.

Step-by-step solution

Step 1 of 5

Ingestion builds the index

Retrieval quality is decided before the user ever asks a question. Ingestion is an offline pipeline that turns an organization's messy sources — PDFs, HTML pages, support tickets, database rows — into a searchable index. Parsing extracts text and preserves structure as metadata. Chunking splits the text into passages small enough to embed meaningfully and large enough to stay self-contained. Each chunk is then embedded by the same model the query path will use, because retrieval only works if both sides live in the same vector space. Finally, vectors plus metadata are upserted into the vector database, idempotently keyed by document id and chunk index so a re-run replaces rather than duplicates. Watch the animation: every phase turns raw representation into something more retrievable. Treat ingestion as a versioned, repeatable job rather than a one-off script, because a stale or partially updated index silently degrades every query that follows.

Animation — Ingestion builds the index

Raw Documents

PDF, HTML, tickets

Parse + Clean

text + structure

Chunk

windows + overlap

Embedding Model

batched

Vector DB Upsert

vectors + metadata

1/6

Documents arrive as PDFs, HTML, tickets, and database rows.

Edge cases & traps

  • Embedding queries with a different model or normalization than ingestion: the query and documents land in different vector spaces, so freeze the embedding model and version the index alongside it.
  • Fixed-size character chunking slices through tables, code, and sentences: use structure-aware splitting (headings, code fences) with 10–20% overlap so boundary facts stay whole.
  • Optimizing only top-k similarity and skipping reranking: the answer sits buried under near-duplicates — retrieve k = 50 fast, then rerank to n = 5 precise.
  • Hallucination when retrieval misses: give the model an explicit "insufficient context" escape hatch and evaluate retrieval recall separately from answer quality.
  • Stale index: a source is updated but old chunks remain — key chunks by document version, re-ingest on change, and delete superseded vectors.

Follow-up questions

Go deeper: Explore the AI visualizer