Ch 8 Retrievers for Question-AnsweringDesign
Fine-tuning the model that decides which passages the LLM ever sees.
Core concepts
- The ODQA trio. An open-domain question-answering pipeline relies on three main components:
- Vector database. Indexing and storing context representations.
- Retriever model. Mapping incoming queries and source documents into a shared high-dimensional vector space.
- Reader model. Extracting precise answer spans from retrieved candidate contexts.
- The “garbage in, garbage out” constraint. The retriever is the primary bottleneck of the pipeline. If it fails to supply semantically relevant context, downstream reader or LLM components cannot recover correct answers.
- The long-tail decision.
- Short-head domain (general knowledge). Pretrained embedding models perform sufficiently for mainstream queries.
- Long-tail domain (niche data). Custom technical, medical, or legal applications require fine-tuning dense retriever models to bridge vocabulary and semantic gaps.
What you must master
- Pretrained vs. fine-tuned — common domains: pretrained models suffice. Niche domains (“long tail”): fine-tune, since no generic benchmark covers them. Level 1
- Pooling — transformers output per-token vectors; mean-pooling averages them into one sentence vector. This is what turns a language model into an embedding model. Level 2
- Batch construction — MNR loss needs duplicate-free batches. Naive batching leaks duplicate positives in as false negatives. Level 2
- Evaluation — use IR metrics (mAP@K), not classification error. Level 3
- Exhaustive vs. approximate search — comparing against every vector doesn’t scale; approximate search trades some accuracy for speed. Level 3