Ch 12 Unsupervised Training with Query Generation — GenQApply

Inventing the queries you don't have by asking a generative model.

Core concepts

  • The idea. You have passages but no queries. Use a query-generation model (e.g. T5) to synthesize a plausible question for each passage.
  • Synthetic pairs. Each (generated query, passage) becomes a positive training pair — fed to MNR loss to fine-tune the retriever.
  • Asymmetric adaptation. Purpose-built for the query↔passage asymmetry of real search.
  • Noise tolerance. Generated queries are imperfect, but in-batch negatives make training robust to some noise.
GenQ Synthetic Query Pipeline & Asymmetric RetrievalGen AI Solution Architect Reference Architecture • Unsupervised Bi-Encoder Fine-TuningTRACK: SEARCH & RETRIEVALPHASE 1: OFFLINE SYNTHETIC DATA & MODEL FINE-TUNING (GenQ)In-Domain Contexts(Raw Text Passages)Passage P₁: “…”Passage P₂: “…”Passage P₃: “…”Seq2Seq GeneratorT5 / Query-Gen Model• Asymmetric Prompting• Nucleus Sampling (top_p)• Generates N queries/passageOutputs (Q_syn, P_raw) PairsNoDuplicates BatchingBatch Integrity LayerGoal: Avoid In-Batch Collisions• Ensures Pᵢ ≠ Pⱼ for all i ≠ j• Critical for Contrastive LossPrevents MNR Loss DistortionBi-Encoder Fine-TuningSentence TransformerLoss: Multiple Negatives RankingSim(Qᵢ, Pᵢ) >> Sim(Qᵢ, Pⱼ)Maximizes Diagonal Cosine SimOutput: Domain-Adapted ModelDeploy Fine-Tuned ModelPHASE 2: ONLINE LOW-LATENCY INFERENCE & ASYMMETRIC RETRIEVALUser App / QueryAsymmetric SearchShort Query:“Who are Normans?”Latency Budget: <10msFine-Tuned Encoder(GenQ MPNet / Dense)Embed Query → Vector[0.012, -0.231, … 768d]Domain-Aligned SpaceVector DB (Pinecone)ANN Index / Similarity SearchCosine Similarity / HNSW IndexTop 1Fetches Top-K Context MatchesDownstream LLM / RAGGrounded GenerationPrompt Payload:Context: “The Normans were…”Query: “Who are Normans?”Accurate & Grounded OutputARCHITECT’S CHEATSHEET & DESIGN TRADEOFFSWhen to Choose GenQ• Zero Labeled Data: Unlabeled domain text only.• Asymmetric Search: Short queries to long passages.• Low Latency Requirement: Needs fast Bi-encoder.• Cold Start RAG: Bootstrapping domain retrieval.Key Risks & Mitigation• Synthetic Noise: T5 hallucinated query generation.• Batch Duplicates: Breaks MNR loss assumptions.• Vocabulary Drift: Niche domain jargon missing in T5.• Fix: Upgrade to GPL (Cross-Encoder pseudo-labeling).Production Sizing & Stack• Query Gen Model: BeIR/query-gen-msmarco-t5-large• Bi-Encoder Base: MPNet / BGE / E5-base• Vector Engine: Pinecone (HNSW Index)• Training Overhead: One-time offline GPU batch compute.

What you must master

  • Explain the passage→synthetic-query→pair pipeline Level 1
  • Generate synthetic queries and fine-tune a retriever with them Level 2
  • Assess synthetic-query quality and its effect on retrieval Level 2

Architect’s lens

GenQ is often the fastest route to a domain-adapted retriever when you have documents and an LLM but no query logs. It’s the embedding-world version of synthetic-data bootstrapping — a technique you’ll propose whenever labelled interaction data is missing at project kickoff.