by Haytham ElFadeel - hfadeelm@gmail.com, Stan Peshterliev
2021
Research done while @ Meta Inc.
Abstract
Knowledge distillation (KD) transfers information from a teacher model (or an ensemble) to a student model. In addition to compressing large teachers into smaller deployable students, KD is frequently used in a same-capacity regime to improve a single model by distilling “dark knowledge” from an .
This paper formalizes same-size logit distillation with temperature scaling, discusses why naïve same-size KD is often teacher-bounded, and introduces Reliability-Weighted Knowledge Distillation (WKD): a simple per-example weighting scheme that downweights teachers that are incorrect on that example while renormalizing correct teachers to preserve logit scale. The method is evaluated on an internal experiment using an Geoffrey Hinton-style distillation objective and shows improved SQuAD v2.0 F1/EM relative to standard KD in this setting.
1. Introduction
Supervised classification typically trains with one-hot labels. While effective, one-hot targets provide no graded similarity structure among incorrect classes. KD addresses this by training the student to match a teacher’s full predictive distribution, which implicitly encodes class similarity and uncertainty (“dark knowledge”).
1.1 Common KD regimes
KD is commonly used in two regimes:
- Compression KD (large → small): improve a smaller student by transferring teacher knowledge.
- Same-size KD (same capacity): improve a student of comparable size by distilling from a strong single teacher or an ensemble.
This paper focuses on (2): same-size distillation from an ensemble.
2. Preliminaries: Temperature scaling and logit distillation
Let the student produce logits over K classes and teachers produce logits for teacher .
2.1 Temperature-softened distributions
Given logits z, the temperature-softened softmax is:
.
- T=1 recovers the standard softmax.
- T>1 yields a higher-entropy (“softer”) distribution, typically exposing more informative mass over non-argmax classes, which can regularize training.
2.2 Standard KD objective (same-size)
A common same-size KD objective combines:
- a distillation loss between teacher and student softened distributions, and
- a supervised loss against ground-truth labels.
Let be the teacher distribution and the student distribution at temperature T. A standard distillation term is the KL divergence:
.
The overall objective is a weighted sum:
where is typically cross-entropy with T=1, and trades off imitation vs. supervision.
2.3 Ensemble teachers
In same-size KD from an ensemble, teachers are trained first, then their predictions are aggregated (e.g., averaging logits) to form a single soft target for student training.
3. Why same-size KD can be teacher-bounded
Empirically, distilling from a strong teacher or ensemble often improves a student relative to training from one-hot labels, but may not match the teacher/ensemble.
A key issue is teacher error propagation: if the student is optimized to match teacher probabilities (including systematic teacher mistakes), the student can inherit those mistakes, especially when α\alphaα is large or when the teacher distribution is overconfident.
3.1 Attempts to surpass the teacher
Two explored strategies were:
- Teacher annealing: train with stronger teacher imitation early, then gradually emphasize ground-truth labels late.
- Iterative KD / Born-Again Networks: repeatedly train a student, then promote it to teacher for the next generation.
In the reported experiments, neither surpassed the standard KD baseline in this setup.
4. Reliability-Weighted Knowledge Distillation (WKD)
4.1 Motivation
The value of KD comes from enriched targets (dark knowledge), but teachers can be wrong on a non-trivial fraction of examples. Rather than only scheduling when to imitate teachers (annealing), WKD changes what is imitated by downweighting incorrect teachers per example while preserving an overall logit scale.
4.2 Method
Assume teachers. For a training example , teacher outputs logits .
Choose a hyperparameter (e.g., ) to downscale incorrect teachers, and which used to scale correct teachers, is defined to enforce a unit average scaling across teachers.
.
Now define weighted teacher logits:
Aggregate teachers by averaging the weighted logits:
.
The student is trained with the same KD objective using from WKD.
4.3 Discussion
WKD can be interpreted as a per-example teacher gating mechanism:
- Incorrect teachers are suppressed to reduce the influence of noisy targets.
- Correct teachers are amplified to keep the ensemble target in a comparable logit scale regime (avoiding trivial shrinkage of ).
This is a minimal intervention (one scalar W) that targets a concrete failure mode: learning from teacher mistakes.
5. Evaluation
5.1 Setup
KD is applied to an ELECTRA-Large model with multi-task learning (MTL) pretraining and evaluated on SQuAD v2.0.
The compared approaches include:
- No KD baseline.
- An ensemble of 3 models.
- Standard KD from the 3-model ensemble.
- Annealing KD.
- WKD with W=0.75.
5.2 Results
Approach | F1 | EM |
Single model (no KD) | 90.87 | 88.34 |
Ensemble of 3 models | 91.41 | 88.92 |
Standard KD from 3 models | 91.38 | 88.84 |
Annealing KD from 3 models | 91.0 | 88.1 |
WKD (W=0.75) from 3 models | 91.5 | 89.0 |
In this experiment, WKD improves over standard KD and slightly exceeds the ensemble score on both F1 and EM.
6. Limitations and future work
- Single-task evidence: results are reported for one downstream task/model combination (SQuAD v2.0, ELECTRA-Large). Broader validation (GLUE, NLI, other QA datasets, vision benchmarks) would establish robustness.
- Teacher correctness signal: the current gating uses top-1 correctness, which is label-dependent and discrete. Extensions could use calibrated teacher confidence, margin, or verifier-based correctness for structured outputs.
- Hyperparameter sensitivity: WKD introduces W and inherits α\alphaα, TTT. Mapping performance across these knobs (and teacher diversity levels) would be useful.
- Theoretical characterization: WKD is plausibly reducing label noise in the teacher signal; connecting it to noise-robust distillation or mixture-of-experts gating would strengthen the argument.
7. Conclusion
Same-size KD can be an effective way to compress ensemble behavior into a single deployable model, but it is often limited by teacher mistakes. WKD addresses this by reweighting teacher logits per example based on correctness while preserving overall logit scale. In the reported SQuAD v2.0 experiment, WKD improves upon standard KD and yields the best F1/EM among tested variants.
References
- G. Hinton, O. Vinyals, J. Dean. Distilling the Knowledge in a Neural Network. arXiv:1503.02531.
- C. Buciluă, R. Caruana, A. Niculescu-Mizil. Model Compression. KDD 2006. DOI: 10.1145/1150402.1150464.
- T. Furlanello, Z. C. Lipton, M. Tschannen, L. Itti, A. Anandkumar. Born-Again Neural Networks. arXiv:1805.04770.
- K. Clark, M.-T. Luong, Q. V. Le, C. D. Manning. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. arXiv:2003.10555.
- P. Rajpurkar, R. Jia, P. Liang. Know What You Don’t Know: Unanswerable Questions for SQuAD (SQuAD 2.0). arXiv:1806.03822.
- A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, Y. Bengio. FitNets: Hints for Thin Deep Nets. arXiv:1412.6550.
- Q. Xie, M.-T. Luong, E. Hovy, Q. V. Le. Self-training with Noisy Student improves ImageNet classification. arXiv:1911.04252.