Ch 10 Data Augmentation with BERT — AugSBERTApply

Using a slow, accurate cross-encoder to label data for a fast bi-encoder.

Core concepts

  • The scarcity problem. Small labelled datasets give weak bi-encoders.
  • AugSBERT recipe. Train a cross-encoder on the small labelled set → use it to label many new sentence pairs (silver data) → train the bi-encoder on the enlarged set.
  • Pair sampling. Which unlabeled pairs to score matters (random, BM25, kNN sampling) to avoid a flood of trivial negatives.
  • Synthetic Knowledge Distillation (AugSBERT)> Small high-quality “gold” datasets can be leveraged to label massive synthetically generated “silver” datasets using a high-accuracy Cross-Encoder.
AugSBERT Data Augmentation ArchitectureSynthetic Gold-to-Silver Knowledge Distillation PipelineGOLD DATAOriginal Dataset1k–5k Labeled PairsCROSS-ENCODER1. Fine-Tune BERT CEHigh accuracy on small dataGENERATION2. Pair SamplingRandom sampling / nlpaugAUTO-LABELING3. Cross-Encoder ScoringPredicts similarity for pairsSILVER DATAAugmented Dataset100k+ Soft-Labeled PairsTARGET BI-ENCODER4. Fine-Tune SBERT ModelTrained on Gold (Emerald) + Silver (Blue) Data
AugSBERT data-augmentation strategy visualizes the dual-encoder setup, gold-to-silver data transformation, and Cross-Encoder synthetic labeling flow.

What you must master

  • Explain the cross-encoder-as-labeler (“silver data”) idea Level 1
  • Encoder Architecture Trade-offs Level 1
  • Cross-Encoders: Pass query and document simultaneously through full self-attention, delivering peak accuracy at a heavy O(N) inference cost.
    Bi-Encoders: Encode text independently into dense vectors, trading full cross-attention context for sub-second vector search performance.
  • Implement the AugSBERT train→label→train loop Level 2
  • Choose a pair-sampling strategy that yields informative pairs Level 2

Architect’s lens

A pragmatic pattern: spend a strong-but-slow model’s accuracy once, offline, to cheaply expand training data for a fast model you’ll serve at scale. This “distill accuracy into speed” mindset recurs across Gen AI system design (e.g. using a big LLM to generate fine-tuning data for a small one).