Ch 7 Introduction to Open-Domain QADesign

The end-to-end blueprint for answering questions over a large knowledge base.

Core concepts

  • Open-domain QA (ODQA). Answer questions where the answer lives somewhere in a large corpus, not a single provided passage.
  • Retriever → Reader. A retriever finds the most relevant passages (vector search); a reader extracts or generates the answer from them.
  • Three components. Vector database (stores context embeddings), retriever (embeds query, fetches top-K), reader (produces the answer).
  • The direct ancestor of RAG. Swap an extractive reader for a generative LLM and you have Retrieval-Augmented Generation.
QuestionRetrieverVector DBtop-K passagesReader / LLMAnswer
The retriever–reader pipeline. Replace the reader with a generative LLM and you have RAG.

Open-Domain vs. Closed-Domain QA

  • Before the pipeline details, pin down what “domain” means in QA. The domain is how broad and how bounded the set of answerable questions is — it shapes which architecture you reach for, and it is orthogonal to (independent of) how the model finds the answer.
  • Open-domain QA. The question can be about anything — no pre-declared subject area — so the answer may sit anywhere in a large, largely-unbounded knowledge source (or inside the model’s own parameters). This is the setup this whole chapter is about: you can’t enumerate the candidates, so you must retrieve to locate the answer.
  • Closed-domain QA. The question is restricted to a single, curated domain — medicine, law, one company’s product docs, a game’s rulebook. The system only ever sees a bounded, in-scope corpus, and anything outside it is a deliberate “I don’t know,” not a guess.
Open-domain QAquestion can be about anything“Who won the 2022 FIFA World Cup?”sporthistorymedicinetechfoodhuge, largely-unbounded corpusmust retrieve to locate the answerClosed-domain QAquestion restricted to ONE domain“Is metformin first-line for Type 2 diabetes?”medicine  (only)one curated, in-scope corpusout-of-scope → “I don’t know”
Open-domain: the answer can hide in any of many domains, so you search a large corpus. Closed-domain: one curated corpus, everything else is rejected by design.

Within Open-Domain: Open-Book vs. Closed-Book

  • Both open-domain QA systems search a large space for an answer, but they differ in where they look at inference time. The “book” is the body of knowledge the model is allowed to consult while answering:
  • Closed-book QA. The model answers from its own parameters (parametric memory) — the “book is closed” — with no retrieval at inference. Fast and cheap, but only as knowledgeable and current as the training data; it can hallucinate and gives no citations.
  • Open-book QA. The model is handed the most relevant passages from a live corpus at inference — the “book is open.” This is exactly the retriever → reader / RAG pipeline from the top of the chapter: grounded, up-to-date, and citable, at the price of extra latency and infrastructure.
Closed-book QAanswer from the model’s parameters — no retrieval at inferenceQuestionLLM / parameter memorythe “closed book”Answerfast · cheapstale factsno citationsOpen-book QA  (RAG)answer from a live corpus — retrieval happens at inferenceQuestionRetrieveropen corpustop-K passagesReader / LLMAnswergrounded · up-to-date · citable
Closed-book answers straight from memory; open-book retrieves the top-K passages and reads them — the RAG pipeline from the top of this chapter.

Example use cases

  • Closed-domain QA — a hospital formulary assistant. “Is metformin covered under our Pharmacy Tier 2 plan?” Only consults the approved formulary and coverage documents; an out-of-scope drug returns a confident “I don’t know” instead of a guess. The bounded corpus is a feature: it makes accidental, dangerous answers structurally impossible.
  • Open-book QA — a release-notes / tech-doc assistant. “What changed in the latest API v2 release?” Retrieves the newest docs at query time and cites the passages, so it stays correct even after the model’s training cutoff — the whole point of RAG over a live corpus.
  • Closed-book QA — a general-purpose chatbot / trivia. “Who is the current Prime Minister of the UK?” Answered instantly from parametric memory with zero retrieval infrastructure and near-zero latency. Great for common, slow-changing knowledge where a hallucinated fact is low-stakes.

Design note: production open-domain systems usually mix the two — the LLM’s parametric memory supplies common-sense that no corpus documents, while the retriever “opens the book” to ground facts that must be current or verifiable. The architecture question is which facts can decay or be wrong in memory, and therefore must be pulled from the corpus instead.

What you must master

  • Draw the full ODQA pipeline and name each component’s job Level 1
  • Explain how ODQA maps onto modern RAG Level 2
  • Reason about where quality is lost (retrieval miss vs. reader error) Level 3
  • Design the pipeline for a given corpus size, latency & accuracy target Level 3

Architect’s lens

This is the reference architecture you will draw on whiteboards constantly. The critical insight for design and debugging: most RAG failures are retrieval failures — if the right passage never reaches the reader/LLM, no amount of prompt engineering saves you. Instrument retrieval quality first.