Ch 8 Retrievers for Question-AnsweringDesign

Fine-tuning the model that decides which passages the LLM ever sees.

Core concepts

  • The ODQA trio. An open-domain question-answering pipeline relies on three main components:
    • Vector database. Indexing and storing context representations.
    • Retriever model. Mapping incoming queries and source documents into a shared high-dimensional vector space.
    • Reader model. Extracting precise answer spans from retrieved candidate contexts.
  • The “garbage in, garbage out” constraint. The retriever is the primary bottleneck of the pipeline. If it fails to supply semantically relevant context, downstream reader or LLM components cannot recover correct answers.
  • The long-tail decision.
    • Short-head domain (general knowledge). Pretrained embedding models perform sufficiently for mainstream queries.
    • Long-tail domain (niche data). Custom technical, medical, or legal applications require fine-tuning dense retriever models to bridge vocabulary and semantic gaps.

What you must master

  • Pretrained vs. fine-tuned — common domains: pretrained models suffice. Niche domains (“long tail”): fine-tune, since no generic benchmark covers them. Level 1
  • Pooling — transformers output per-token vectors; mean-pooling averages them into one sentence vector. This is what turns a language model into an embedding model. Level 2
  • Batch construction — MNR loss needs duplicate-free batches. Naive batching leaks duplicate positives in as false negatives. Level 2
  • Evaluation — use IR metrics (mAP@K), not classification error. Level 3
  • Exhaustive vs. approximate search — comparing against every vector doesn’t scale; approximate search trades some accuracy for speed. Level 3

Architect’s lens

  • Encoding timing = latency budget. Contexts encode offline, in batch. Queries encode on every request. Keep re-embedding off the request path.
  • Retriever and reader fail independently. A retriever regression silently corrupts everything downstream — the reader will confidently answer from junk. Monitor retrieval quality separately from end-to-end accuracy.
  • Domain fit beats model size. A model fine-tuned on your data will likely beat a stronger general-purpose model in-domain. Optimize for data fit, not leaderboard rank.
  • Metadata is cheap now, expensive later. Filtering by category, tenant, or freshness is far easier to design in from the start than retrofit.
  • Approximate search is a dial, not a fixed setting. Recall vs. latency is a product decision — tune it to how much retrieval error the reader can tolerate.