Fetching the paper…
Reading the bibliography…
Ensemble knowledge distillation can extract knowledge from multiple teacher models and encode it into a single student model.
Random decision forests. In ICDAR , Vol. 1. IEEE, 278–282
Tin Kam Ho. 1995 · 1995
Earlier work this paper cites.
A brief introduction to boosting. In IJCAI , Vol. 99. 1401–1406
Robert E Schapire. 1999 · 1999
Earlier work this paper cites.
Ensemble methods in machine learning. In International workshop on multiple classifier systems . Springer, 1–15
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
Random k-labelsets: An ensemble method for multilabel classification. In European conference on machine learning . Springer, 406–417
Grigorios Tsoumakas and Ioannis Vlahavas. 2007 · 2007
Earlier work this paper cites.
Selective sampling and active learning from single and multiple teachers
Ofer Dekel, Claudio Gentile, and Karthik Sridharan. 2012 · 2012
Earlier work this paper cites.
Ensemble approaches for regression: A survey
Joao Mendes-Moreira, Carlos Soares, Alípio Mário Jorge, and Jorge Freire De Sousa. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP . 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In ICLR
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Distilling Knowledge from Ensembles of Neural Networks for Speech Recognition.. In Interspeech . 3439–3443
Yevgen Chebotar and Austin Waters. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Ensemble distillation for neural machine translation
Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran. 2017 · 2017
Earlier work this paper cites.
Efficient Knowledge Distillation from an Ensemble of Teachers.. In Interspeech . 3697–3701
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, and Bhuvana Ramabhadran. 2017 · 2017
Earlier work this paper cites.
Data distillation: Towards omni-supervised learning. In CVPR . 4119–4128
Ilija Radosavovic, Piotr Dollár, Ross Girshick, Georgia Gkioxari, and Kaiming He. 2018 · 2018
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In BlackboxNLP . 353–355
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL . 1112–1122
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Xiatian Zhu, Shaogang Gong, et al · 2018
Cited alongside, same era.
On the efficacy of knowledge distillation. In CVPR . 4794–4802
Jang Hyun Cho and Bharath Hariharan. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Agree to disagree: Adaptive ensemble knowledge distillation in gradient space
Shangchen Du, Shan You, Xiaojie Li, Jianlong Wu, Fei Wang, Chen Qian, and Changshui Zhang. 2020 · 2020
Later among the works it cites.
Ensemble learning of lightweight deep learning models using knowledge distillation for image classification
Jaeyong Kang and Jeonghwan Gwak. 2020 · 2020
Later among the works it cites.
Improving model calibration with accuracy versus uncertainty optimization
Ranganath Krishnan and Omesh Tickoo. 2020 · 2020
Later among the works it cites.
Ensemble distillation for robust model fusion in federated learning
Tao Lin, Lingjing Kong, Sebastian U Stich, and Martin Jaggi. 2020 · 2020
Later among the works it cites.
Feded: Federated learning via ensemble distillation for medical relation extraction. In EMNLP . 2118–2128
Dianbo Sui, Yubo Chen, Jun Zhao, Yantao Jia, Yuantao Xie, and Weijian Sun. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daliang Li and Junpu Wang. 2019 · 2019
Cited alongside, same era.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019a · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 2019
Cited alongside, same era.
Feed: Feature-level ensemble for knowledge distillation
SeongUk Park and Nojun Kwak. 2019 · 2019
Cited alongside, same era.
Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion. In Interspeech . 2115–2119
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2020 · 2020
Cited alongside, same era.
Unilmv2: Pseudo-masked language models for unified language model pre-training. In ICML . PMLR, 642–652
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al · 2020
Cited alongside, same era.
Online ensemble model compression using knowledge distillation. In ECCV . Springer, 18–35
Devesh Walawalkar, Zhiqiang Shen, and Marios Savvides. 2020 · 2020
Later among the works it cites.
Distilling knowledge from an ensemble of convolutional neural networks for seismic fault detection
Zirui Wang, Bo Li, Naihao Liu, Bangyu Wu, and Xu Zhu. 2020 · 2020
Later among the works it cites.
Improving bert fine-tuning via self-ensemble and self-distillation
Yige Xu, Xipeng Qiu, Ligao Zhou, and Xuanjing Huang. 2020 · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization. In CVPR . 3903–3911
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng. 2020 · 2020
Later among the works it cites.
Ensemble Attention Distillation for Privacy-Preserving Federated Learning. In ICCV . 15076–15086
Xuan Gong, Abhishek Sharma, Srikrishna Karanam, Ziyan Wu, Terrence Chen, David Doermann, and Arun Innanje. 2021 · 2021
Later among the works it cites.
Specific Expert Learning: Enriching Ensemble Diversity via Knowledge Distillation
Wei-Cheng Kao, Hong-Xia Xie, Chih-Yang Lin, and Wen-Huang Cheng. 2021 · 2021
Later among the works it cites.
One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers. In ACL Findings . 4408–4413
Chuhan Wu, Fangzhao Wu, and Yongfeng Huang. 2021 · 2021
Later among the works it cites.