Ch 10 Data Augmentation with BERT — AugSBERTApply
Using a slow, accurate cross-encoder to label data for a fast bi-encoder.
Core concepts
- The scarcity problem. Small labelled datasets give weak bi-encoders.
- AugSBERT recipe. Train a cross-encoder on the small labelled set → use it to label many new sentence pairs (silver data) → train the bi-encoder on the enlarged set.
- Pair sampling. Which unlabeled pairs to score matters (random, BM25, kNN sampling) to avoid a flood of trivial negatives.
- Synthetic Knowledge Distillation (AugSBERT)> Small high-quality “gold” datasets can be leveraged to label massive synthetically generated “silver” datasets using a high-accuracy Cross-Encoder.
AugSBERT data-augmentation strategy visualizes the dual-encoder setup, gold-to-silver data transformation, and Cross-Encoder synthetic labeling flow.
What you must master
- Explain the cross-encoder-as-labeler (“silver data”) idea Level 1
- Encoder Architecture Trade-offs Level 1 Cross-Encoders: Pass query and document simultaneously through full self-attention, delivering peak accuracy at a heavy O(N) inference cost.
- Implement the AugSBERT train→label→train loop Level 2
- Choose a pair-sampling strategy that yields informative pairs Level 2
Bi-Encoders: Encode text independently into dense vectors, trading full cross-attention context for sub-second vector search performance.
Architect’s lens
A pragmatic pattern: spend a strong-but-slow model’s accuracy once, offline, to cheaply expand training data for a fast model you’ll serve at scale. This “distill accuracy into speed” mindset recurs across Gen AI system design (e.g. using a big LLM to generate fine-tuning data for a small one).