GPL Architecture Pipeline

Offline & Online

Generative Pseudo-Labeling for domain adaptation of dense retrievers (SBERT).

docs synth triples score loss deploy Raw Corpus Unlabeled Docs 1. Query Gen T5 / doc2query 2. Neg Miner Dense Retrieval 3. Teacher Cross-Encoder 4. Student Tr. Margin MSE Loss Fine-Tuned Bi-Enc SBERT Model Runtime Vector Index (Pinecone) Online Serving Infrastructure (<15ms SLA)
Offline Pre-training / Pseudolabeling
Fine-Tuned Model & Serving

Hover over any node

INSPECTION
Inputs
Select a pipeline step above to inspect inputs & outputs.
Outputs
Displays exact data payloads passed between stages.
Architectural Considerations
GPL achieves state-of-the-art domain adaptation without manual labels by synthesizing queries, mining hard negatives, and distilling cross-encoder margins into a fast bi-encoder.