By Haytham ElFadeel - hfadeelm@gmail.com
2018
Abstract
Label smoothing (LS) is a regularization technique that replaces one-hot training labels with soft targets, typically by mixing the hard label with a uniform distribution. The practical effect is to reduce pathological over-confidence (very low-entropy predictive distributions), which often improves generalization and calibration.
However, uniform smoothing implicitly assumes that all incorrect classes are equally plausible. In sequence problems—especially tasks where the label corresponds to a token position (e.g., extractive QA start/end indices) or a boundary—this assumption is often too coarse: near-miss predictions (off by ±1–2 tokens) are typically more plausible than far-away ones.
Gaussian Label Smoothing (GLS) replaces the uniform “noise” distribution with a localized Gaussian kernel centered at the gold token position, allocating more probability mass to nearby positions than to distant ones.
1. Introduction
Consider a classification problem with classes. Let be the gold label and the model distribution.
One-hot target:
Label smoothing with (uniform prior):
Training minimizes cross-entropy .
A useful view is that Label Smoothing adds a term that discourages extremely peaked distributions, often improving generalization.
2. Gaussian Label Smoothing (GLS)
2.1. Where GLS is most natural: token-position classification
In extractive QA and related tasks, models often predict a categorical distribution over token indices. For a sequence of length N, the model predicts:
- Start distribution over
- End distribution over
Let be the gold index (start or end). Define a discrete Gaussian kernel over positions:
Then GLS defines the smoothed target:
The loss is standard cross-entropy:
Intuition: GLS explicitly encodes the inductive bias that “off-by-one” boundaries are closer to correct than arbitrarily distant ones—common in span extraction error patterns.
2.2. Practical hyperparameters
- controls how much probability mass is moved away from the hard target (typical values are small, e.g., 0.05, 0.10).
- controls the neighborhood width. A value around 1–2 tokens often matches “boundary jitter” seen in span prediction; larger approaches a flatter distribution (closer in spirit to uniform smoothing).
2.3. Extending GLS beyond span indices
For token-level tagging (e.g., BIO NER), the label space is not inherently ordered like positions. GLS is most principled when you have:
- an ordering/metric over labels, or
- an intermediate formulation where predictions correspond to structured boundaries (e.g., span-based tagging).
The original post motivates GLS from the observation that in language, neighboring tokens and boundaries are correlated, so “nearby” mistakes should be penalized less aggressively than “far” mistakes.
3. Relationship to knowledge distillation
Knowledge distillation (KD) trains a student model using soft targets from a teacher (or teacher ensemble), often written as a mixture of one-hot labels and teacher probabilities.
GLS can be viewed as a hand-crafted proxy for a teacher distribution in boundary-sensitive tasks:
- KD supplies a learned non-uniform distribution (teacher “dark knowledge”).
- GLS supplies a fixed, geometry-aware distribution centered on the gold token, reflecting common near-miss structure (e.g., slightly shifted start/end indices).
4. Experiments
BERT - Base in CoNLL-2003 Named Entity Recognition
System | Dev F1 | Test F1 |
BERT - Base (paper) | 96.4 | 92.4 |
BERT - Base (reproduced) | 96.4 | 92.4 |
BERT - Base + LS | 96.4 | 92.4 |
BERT - Base + GLS (our) | 96.5 | 92.6 |
BERT - Large in SQuAD 2.0
System | Dev F1 | EM |
BERT - Large (reproduced) | 84.1 | 81.0 |
BERT - Large + LS | 84.2 | 81.1 |
BERT - Large + GLS (our) | 84.9 | 81.4 |
ALBERT - XLarge in SQuAD 2.0
System | Dev F1 | EM |
ALBERT - XLarge (reproduced) | 87.9 | 84.1 |
ALBERT - XLarge + LS | 88.1 | 84.5 |
ALBERT - XLarge + GLS (our) | 88.2 | 85.2 |
References
- Haytham ElFadeel. Gaussian Label Smoothing (blog post).
- Christian Szegedy et al. Rethinking the Inception Architecture for Computer Vision (introduces label-smoothing regularization).
- Geoffrey Hinton, Oriol Vinyals, Jeff Dean. Distilling the Knowledge in a Neural Network (knowledge distillation).
- Jacob Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.
- Pranav Rajpurkar, Robin Jia, Percy Liang. Know What You Don’t Know: Unanswerable Questions for SQuAD (SQuAD 2.0).
- Erik F. Tjong Kim Sang, Fien De Meulder. Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.
- Zhenzhong Lan et al. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.
- Rafael Müller, Simon Kornblith, Geoffrey Hinton. When Does Label Smoothing Help?