Skip to content
Reliable Data Engineering
Practice problem hard ragllmvector-searchembeddingsevaluation
Practise with timer, notes and rubric

Design an Enterprise RAG Knowledge Assistant

RAG pipeline

Problem

Employees should be able to ask questions in natural language and get answers grounded in internal knowledge: Confluence, SharePoint, Google Drive, Jira, Slack, PDFs. Answers must cite sources and only use documents the asking user is allowed to see. Design the system with a focus on the data pipelines.


Clarifying questions

QuestionAssumed answer
Corpus size?5 M documents, ~50 M chunks; 2% change per day
Users / QPS?50k employees, peak 50 queries/s
Latency?First token < 2 s, full answer < 8 s
Freshness?Edits visible within 15 minutes; permission changes within 5 minutes
Languages?English + German + Japanese
Hosting constraints?LLM via a cloud provider in-region; no training on company data

1. Estimates

Chunks: 50M × 1024-dim float32 (4 KB) = 200 GB vectors → + HNSW overhead ≈ 300 GB RAM
  → int8 quantisation ≈ 75 GB, or sharded index across nodes
Daily changes: 2% × 5M = 100k docs → ~1M chunks re-embedded/day → ~500M tokens/day → tens of $/day
Initial embedding: 50M chunks × 500 tokens = 25B tokens → one-off, hundreds of $; rate limits dominate time (days)
Query: 50 QPS × (retrieve 50 + rerank + LLM ~4k input tokens) → LLM cost dominates: budget per query

2. Architecture

flowchart LR
    subgraph Connectors
        C1[Confluence] 
        C2[SharePoint / Drive]
        C3[Jira / Slack]
    end
    C1 --> Q[[Change queue<br/>webhooks + periodic crawl]]
    C2 --> Q
    C3 --> Q
    Q --> FETCH[Fetch + raw store<br/>bronze: original files]
    FETCH --> PARSE[Parse, clean, PII redact<br/>layout-aware]
    PARSE --> CHUNK[Chunk + metadata<br/>doc_id, section, acl, lang, updated_at]
    CHUNK --> DIFF{chunk hash changed?}
    DIFF -->|yes| EMB[Embedding workers<br/>batched, rate-limited]
    DIFF -->|no| SKIP[skip]
    EMB --> CT[(silver.chunks: text, vector,<br/>metadata, model_version)]
    CT --> VI[(Vector + BM25 index<br/>synced from table)]
    ACL[ACL sync service<br/>group memberships, doc permissions] --> CT
    ACL --> VI
    U[User + SSO identity] --> API[Assistant API]
    API --> VI
    API --> RR[Reranker]
    API --> LLM[LLM]
    API --> LOG[(llm_ops.requests)]
    LOG --> EVAL[Eval pipelines + dashboards]

3. Deep dives

3.1 Incremental ingestion

3.2 Permissions (the make-or-break requirement)

flowchart LR
    DOCP[Doc permissions<br/>from source APIs] --> NORM[Normalise to principal ids<br/>users + groups]
    GRP[IdP group memberships] --> EXP[(user → groups<br/>cache, 5-min refresh)]
    NORM --> META[chunk.acl = list of principal ids]
    Q[Query from user] --> EXP
    EXP --> FILTER["Retrieval filter:<br/>chunk.acl ∩ user principals ≠ ∅"]
    META --> FILTER

3.3 Retrieval quality

3.4 Evaluation & monitoring

WhatHow
Retrieval recall@kGolden set: 500 questions with known relevant docs (from SMEs + mined from search logs)
Faithfulness / citation correctnessLLM judge + human spot checks weekly
Freshnessnow − source.updated_at for recently edited docs, end-to-end ingestion lag
ACL correctnessSynthetic test users with known permissions; automated leakage tests in CI
Cost & latencyPer request logs; budget alerts
User feedback👍/👎 + “wrong source” flags → triage queue → golden set growth

Every change (chunker, embedding model, prompt, reranker) runs the eval suite before rollout; embedding model changes use blue/green indexes.

4. Trade-offs

DecisionChoiceAlternative
Source of truthDelta chunks table; index derivedVector DB as the only store (hard to rebuild/audit)
ACL enforcementPre-filter in retrievalPost-filter, or separate index per group (explodes)
Embedding modelMultilingual, 768–1024 dims, quantisedLarger dims (cost/RAM), per-language models (complexity)
FreshnessWebhooks + crawl + reconcileNightly full re-index (stale, costly)

5. Failure modes

FailureHandling
Embedding API throttled/outageQueue backlog grows; freshness SLO alert; keep serving existing index
Permission sync lagFail closed: if ACL data for a doc is older than threshold, exclude it
Bad parser release mangles tablesEval regression catches; re-parse from bronze raw files
Prompt injection inside documentsTreat retrieved text as data in the prompt, strip/flag suspicious instructions, restrict tool access

6. What separates a senior answer

7. Follow-up questions

How do you handle questions over structured data ("revenue in Q3 for Germany")?

RAG over documents is the wrong tool. Route to a text-to-SQL tool over governed gold tables (semantic layer / metric definitions as context, executed with the user’s permissions), or to certified dashboards. A router (classifier or the LLM with tools) decides between document search and SQL.

Users complain the bot answers from outdated pages. Fix it.

Measure freshness lag; ensure deleted/archived pages are removed (reconciliation crawl); deduplicate near-identical versions; boost recency and prefer canonical/verified spaces; display “last updated” with citations; let owners mark docs as deprecated (metadata filter).


Self-assessment rubric