RAG and retrieval
Grounding answers in your own data: chunking, embeddings, hybrid retrieval and reranking, the context budget, permissions, freshness and citations.
RAG is the design skill behind trustworthy AI answers over private knowledge. This page gets you ready to build the ingest pipeline, the query-time retriever, the freshness story and the citation contract.
Read this if your last attempt…
- Your answer says “put our docs in a vector database” and stops there
- You are not sure where chunking, embeddings, keyword search and reranking fit in the request path
- You forgot to filter by permissions until after the model saw the retrieved text
- Your RAG design has no plan for stale documents, deletes or citations
The concept
What RAG is
Retrieval-augmented generation means the model answers with context fetched at request time rather than only from its trained parameters. The original RAG paper describes models that combine parametric memory with a retrieved non-parametric memory, and calls out provenance and updating world knowledge as problems this shape helps address (Lewis et al., 2020).
Ingest turns source content into searchable chunks with embeddings, metadata and ACLs. Query time filters by permission, retrieves and reranks evidence, then asks the model to answer with citations.
Retrieval designs: choose the simplest path that preserves relevance, safety and freshness.
| Approach | Best fit | Main risk |
|---|---|---|
| Prompt stuffing | Tiny curated knowledge where the full reference comfortably fits the prompt. | No freshness pipeline and no scalable permissions story. |
| Keyword-only retrieval | Exact terms, SKUs, error codes, legal phrases and navigation queries. | Paraphrases and vague user language miss the right document. |
| Vector-only retrieval | Semantic discovery where exact wording varies and identifiers are not central. | Rare names, codes and permission filters can be weak unless metadata is designed well. |
| Hybrid retrieval | Support, docs, policies and enterprise search where exact terms and meaning both matter. | You must combine result lists and tune candidate windows with judgement data. |
| Hybrid plus reranking | High-value answers where the final few chunks must be in the right order. | Extra latency and model cost in the request path. |
How interviewers grade this
- You draw both paths: offline ingest and online query serving.
- You explain chunking, embeddings and metadata as one unit, not separate chores.
- You use hybrid retrieval when exact identifiers and semantic matches both matter.
- You filter permissions before retrieval and ranking, not after generation.
- You budget context and rerank evidence instead of dumping every candidate into the prompt.
- You state how updates, deletes and permission changes invalidate old chunks.
- You require citations that map claims back to source spans.
Variants
Vector-only RAG
Embed the query, retrieve nearest chunks, and prompt the model with those chunks.
This is the fastest prototype. It works when the corpus is small, language is natural, and exact identifiers are not the main access path. The trade-off is recall: semantic matches can miss part numbers, log lines, customer names or legal terms that keyword search would find. Use it to learn the shape, but be ready to add keyword retrieval and metadata filters.
Pros
- +Simple to build and explain.
- +Good for broad semantic questions.
- +Makes the embedding quality visible early.
Cons
- −Weak for exact identifiers and rare terms.
- −Can hide permission and freshness issues behind a demo.
- −Often needs reranking once the corpus grows.
Choose this variant when
- Prototype a low-risk assistant over a narrow, clean corpus.
Hybrid retrieval with reciprocal rank fusion
Run keyword and vector retrieval, then fuse ranked lists before trimming to the prompt budget.
Hybrid retrieval gives keyword search and vector search separate chances to find evidence. RRF is a common rank-based fusion method because it uses ranks instead of assuming BM25 and vector scores mean the same thing. Use this when the user may ask with vague language but the answer still depends on exact product names, policy clauses or identifiers.
Pros
- +Captures exact terms and semantic paraphrases.
- +RRF avoids fragile score normalisation.
- +Works well with a later reranker.
Cons
- −Two retrieval paths mean more tuning and observability.
- −Candidate windows and duplicate removal matter.
- −Still needs permission and freshness filters before ranking.
Choose this variant when
- Most production RAG over docs, support content, policies or code.
Reranked RAG
Use a second-stage ranker to order the best candidates for the exact question.
The first-stage retriever favours recall. A first-stage bi-encoder embeds the query and chunk separately; a cross-encoder reranker scores each question-passage pair together, which is slower per pair but often better for final ordering (Sentence Transformers CrossEncoder docs). This helps when many chunks look similar, or when the answer depends on a condition hidden in the middle of a passage. The trade-off is latency and another model-like component to monitor.
Pros
- +Improves final context quality when first-stage recall is broad.
- +Can prefer directly answerable passages over loosely related ones.
- +Gives a clean place to add diversity and freshness boosts.
Cons
- −Adds request latency.
- −Costs more than plain retrieval.
- −Needs evaluation data to justify the extra stage.
Choose this variant when
- Important answers, noisy corpora, long documents, or many near-duplicate chunks.
Freshness-first RAG
Optimise the ingest path for updates, deletes and permission changes before fancy ranking.
For internal knowledge, the worst answer is often a confident stale one. A freshness-first design keeps source versions, tombstones and ACL changes close to retrieval. It may choose simpler ranking if that lets delete events block stale chunks immediately. This is common for policy, legal, security and customer-data assistants.
Pros
- +Safer when content changes frequently.
- +Clear delete and permission story.
- +Easier incident response because every chunk has a source version.
Cons
- −More metadata and indexing machinery up front.
- −Backfills and version conflicts need careful runbooks.
- −May limit which vector index layouts are acceptable.
Choose this variant when
- Policies, entitlements, customer records, legal content and any corpus with strict deletion needs.
Worked example
Numbers in this section are illustrative.
Scenario: design a RAG assistant for a SaaS support team. It answers from product docs, internal runbooks and resolved tickets. Numbers in this section are illustrative.
Requirements. Agents ask natural language questions. Answers cite sources, avoid customer data the agent cannot access, reflect deletes or permission changes, and say so when evidence is missing.
Ingest. Suppose the corpus has 80,000 docs and tickets. The pipeline watches document events, extracts text, and chunks by structure: headings, FAQ pairs, and issue plus resolution for tickets. A starting chunk target might be around 500 tokens with a small overlap, but the real unit is answerability. Each chunk stores source id, version, URL, offset, tenant, product area, ACL groups and delete state.
Embedding and indexing. The indexer embeds active chunks and writes vectors, text and metadata to the retrieval store. The source system remains the authority. Updates create a new version and retire old chunk ids. Deletes first write a tombstone keyed by source_doc_id, then remove vectors in the background. Backfills are version-aware so older jobs cannot reintroduce stale chunks.
Query path. The API receives: “How do I recover a failed invoice sync for a customer on EU hosting?” It authenticates the agent, builds tenant, role, product and region filters, then runs hybrid retrieval. Keyword search catches “invoice sync” and “EU hosting”; vector search catches paraphrases such as “billing export failed”. ACL and tombstone filters apply before ranking.
Rerank and context. First-stage retrieval might gather 60 candidates, then a reranker trims to 8 passages that answer the question and cover distinct sources. The prompt builder keeps only needed quoted spans plus citation ids. If two chunks conflict, prefer the newer source version or ask for clarification.
Answer contract. The model is instructed to answer only from supplied passages, cite each operational step, and say “I do not know from the retrieved sources” if evidence is missing. The UI resolves citations to document sections or tickets that the user can access.
Operational checks. Track retrieval misses, zero-result queries, citation click-through, stale-hit incidents and permission-denied attempts. Replay past questions when chunking, embeddings, hybrid weights or reranking changes.
Good vs bad answer
Numbers in this section are illustrative.
Interviewer probe
“Design a RAG system for internal support docs. How do you make the answers useful and safe?”
Weak answer
“I would embed all the documents, store them in a vector database, retrieve the top matches and put them into the prompt. The model can cite the docs it saw.”
Strong answer
“I would draw two paths. Ingest parses source docs, chunks them by document structure, embeds each chunk, and stores text plus metadata: source id, version, URL, offsets, ACL fields and deletion state. Query time authenticates the agent, applies tenant and ACL filters before retrieval, runs hybrid keyword plus vector search, reranks candidates, then fits the best spans into the prompt. The model is instructed to answer only from those spans and cite source versions. Updates re-index changed chunks; deletes write tombstones immediately so the query path can exclude stale chunks while cleanup catches up. If the retrieved evidence does not answer the question, the assistant says it does not know instead of guessing.”
Why it wins: The strong answer treats RAG as a system with ingest, retrieval, permissions, freshness, context budgeting and a citation contract. The weak answer names the store but misses the failure modes that make enterprise RAG unsafe.
When it comes up
- A product asks for a chatbot over internal docs or customer data.
- The prompt says the model should answer from private knowledge or cite sources.
- Search quality, freshness, permissions or hallucinations become the hard part.
- The interviewer asks why a vector database alone is not enough.
Order of reveal
- 11. Draw the two paths. Start with ingest and query lanes. Ingest chunks, embeds and stores metadata; query authenticates, retrieves, reranks and prompts.
- 22. Define chunk metadata. Every chunk carries source id, version, URL, offsets, ACL fields and deletion state because retrieval, freshness and citations all depend on metadata.
- 33. Use hybrid retrieval when exact terms matter. Keyword search catches names and codes; vector search catches paraphrases. Fuse and rerank before spending prompt budget.
- 44. Put permissions in retrieval. The model must not see chunks the user cannot access. ACL filters are predicates on retrieval, not a final display filter.
- 55. Close with freshness and citations. Updates re-index affected chunks, deletes create tombstones, and answers cite source spans. Missing evidence becomes “I do not know”.
Signature phrases
- “RAG has an ingest path and a query path.” — Prevents a database-only answer.
- “The vector is a derived index; the source document remains the authority.” — Sets up freshness, deletes and citations.
- “Filter before retrieval, not after generation.” — Shows the permission boundary.
- “Retrieve broadly, rerank narrowly, cite precisely.” — Summarises context budget and grounding.
Likely follow-ups
?“How do you choose chunk size?”Reveal
Choose by answerability, not by a magic token count. Start from document structure, keep a rule with its exception, keep FAQ question and answer together, and inspect misses. If the right answer is split across chunks, merge or overlap. If chunks contain too many unrelated facts, split smaller.
?“Why not just use vector search?”Reveal
Vector search is strong for paraphrase and weak for rare exact identifiers. Support and policy questions often contain product names, error codes, versions and legal terms. Hybrid retrieval lets keyword and vector candidates both compete, then a reranker decides which passages answer the question.
?“What happens when a document is deleted?”Reveal
The source emits a delete event. The retrieval service records a tombstone immediately and excludes that source id from results. A background job removes vectors and old chunks. If cleanup lags, the tombstone still blocks stale retrieval.
?“How do citations work?”Reveal
Citations are not generated from memory. The prompt builder passes labelled spans with source ids and URLs. The model cites those labels, and the app resolves them to source sections the user can access. Unsupported claims should be omitted or marked unknown.
Code examples
type Principal = {
tenantId: string;
groupIds: string[];
regions: string[];
};
type RetrievalRequest = {
question: string;
principal: Principal;
filters: {
tenantId: string;
allowedGroups: string[];
allowedRegions: string[];
deletedAt: null;
};
};
function buildRetrievalRequest(question: string, principal: Principal): RetrievalRequest {
return {
question,
principal,
filters: {
tenantId: principal.tenantId,
allowedGroups: principal.groupIds,
allowedRegions: principal.regions,
deletedAt: null,
},
};
}CREATE TABLE rag_chunks (
source_doc_id TEXT NOT NULL,
chunk_id TEXT PRIMARY KEY,
source_version TEXT NOT NULL,
chunk_text TEXT NOT NULL,
tenant_id TEXT NOT NULL,
acl_groups TEXT[] NOT NULL,
canonical_url TEXT NOT NULL,
deleted_at TIMESTAMPTZ
);
CREATE TABLE rag_tombstones (
source_doc_id TEXT PRIMARY KEY,
deleted_at TIMESTAMPTZ NOT NULL,
reason TEXT NOT NULL
);Common mistakes
Arbitrary splits can separate a rule from its exception or a question from its answer. Start from the source shape: headings, FAQ pairs, tickets, code symbols or policy sections. Then tune size and overlap from retrieval failures.
Embeddings are good at meaning, not guaranteed exact matching. Error codes, SKUs, customer names and legal phrases often need keyword search. Hybrid retrieval gives exact terms and semantic matches separate paths into the candidate set.
If the model saw a confidential chunk, hiding the citation later does not undo the leak. Tenant, group, ownership, region and deletion filters belong in the retrieval query or index layout before ranking begins.
A vector index is derived state. Source updates must retire old chunks, and deletes need tombstones that the query path checks immediately. Background cleanup is fine only after the tombstone blocks stale hits.
More passages can make the answer worse when they dilute the best evidence. Retrieve broadly, rerank, remove duplicates and trim to the smallest set that answers the question with citations.
A citation to a whole wiki or a search result page is hard to verify. Carry source id, version and span offsets from ingest so the answer can cite the exact passage that supports each claim.
Retrieved content is data. A malicious or compromised document can contain instructions aimed at the model. Delimit retrieved passages, keep tool permissions outside the model, and do not let source text override system instructions.
Practice drills
Numbers in this section are illustrative.
A RAG answer cites the right wiki but the wrong paragraph. What broke?Reveal
The citation contract is too coarse. Store span offsets or section anchors with each chunk, pass labelled spans to the prompt, and render citations to the exact source version. Whole-document citations are hard to verify and hide retrieval mistakes.
A deleted HR policy still appears in answers for a few minutes. What is the fix?Reveal
Do not wait for vector cleanup to finish. Record a tombstone from the delete event and make the query path exclude that source id immediately. Then remove the old vectors in the background and verify the tombstone disappears only when cleanup is complete.
Why might hybrid retrieval beat vector-only retrieval for support docs?Reveal
Support questions often contain exact names, error codes and versions. Keyword search is good at those. Vector search catches paraphrases. Hybrid retrieval lets both result sets contribute before reranking chooses the passages that directly answer the question.
Where should access control happen in the RAG path?Reveal
Before the model sees content. Apply tenant, group, ownership, region and delete filters in retrieval or in the index layout. A post-generation filter can hide the citation, but it cannot prevent the model from using confidential text it already saw.
Deep dives
Chunking playbook
Chunking is the most interviewable part of RAG because it looks simple and fails quietly. The goal is not equal-sized text blocks. The goal is evidence units: passages that can answer a specific kind of question and carry enough metadata to be checked later.
Start by listing source shapes. Product docs have headings and subsections. FAQs have a question, accepted answer and maybe examples. Tickets have symptoms, diagnostics and resolution. Code has symbols and comments. Policies have rules, exceptions and effective dates. Each shape suggests a chunk boundary that a blind token splitter would miss.
Then test for answerability. Take real or illustrative questions and ask: would one chunk contain the answer, the condition and the citation? If the answer crosses many chunks, the chunks are too small or the source needs a parent summary chunk. If one chunk answers several unrelated questions, it is too broad and wastes context budget.
Metadata is part of chunking. A chunk without a stable source id, version, title, URL and offset cannot produce reliable citations. A chunk without tenant and ACL fields cannot be safely retrieved. A chunk without updated time and deletion state cannot support freshness. The embedding is only one field in the row.
Use overlap sparingly. Overlap helps when boundaries cut a sentence or rule, but large overlap creates duplicate candidates that waste reranker effort. A diversity step can collapse near-duplicate chunks from the same source before prompt construction.
Finally, maintain chunks as derived data. When the source changes, the old chunks are retired and new chunks get new ids or versions. If chunk ids are unstable, citations and feedback loops break. If old chunks linger, the assistant can answer from stale text even when the source is fixed.
Permission-safe retrieval
Permission-safe RAG starts by treating retrieval as a data access operation, not as search decoration. The user principal becomes part of the retrieval request. Tenant, document state, group membership, ownership and geography should narrow the corpus before ranking decides relevance.
The risky shortcut is retrieve-everything-then-filter. It can leak in several ways. The model may use a confidential passage even if the UI hides the citation. A reranker may prefer private evidence and push public evidence out of the prompt. Logs and traces may store disallowed chunks. Even if no leak is visible, your relevance metrics become misleading because the best public answer may not be selected.
There are two common layouts. In a shared index, ACL fields are indexed metadata and every query carries filters. This is simpler but depends on the engine applying filters early enough for your risk. In a partitioned layout, high-risk tenants or regions get separate indexes, tables or partitions, so a query cannot search outside its boundary. This costs more operations but gives a stronger blast-radius story.
The interview answer should mention group changes too. When a user loses access, cached retrieval results and prompt traces need expiry. When a document changes ACL, the index must update metadata or tombstone old chunks. Permission safety is a freshness problem as much as an auth problem.
Close with observability: log why a chunk was eligible without logging sensitive text by default. Count denied candidates, empty authorised results and tombstone hits. Those metrics catch broken filters before a user sees the wrong answer.
Cheat sheet
- •RAG = ingest path plus query path, not just a vector database.
- •Chunk by answerable units and store source id, version, offsets, ACL and delete state.
- •Use keyword for exact terms, vector for meaning, reranker for final order.
- •Filter permissions before retrieval or inside the index layout.
- •Treat vectors as derived state; source systems own truth.
- •Deletes need tombstones before background cleanup.
- •Context budget is a funnel: retrieve broad, rerank, dedupe, trim.
- •Grounding means claims cite source spans and unsupported answers say unknown.
- •Retrieved text is untrusted data, not an instruction channel.
Practice this skill
These problems exercise RAG and retrieval. Try one now to apply what you just learned.
Read this if