Fetching the paper…
Reading the bibliography…
With ever growing scale of neural models, knowledge distillation (KD) attracts more attention as a prominent tool for neural model compression.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, and Hassan Ghasemzadeh. 2019 · 1902
Earlier work this paper cites.
Essence knowledge distillation for speech recognition
Zhenchuan Yang, Chun Zhang, Weibin Zhang, Jianxiu Jin, and Dongpeng Chen. 2019 · 1906
Earlier work this paper cites.
Bam! born-again multi-task networks for natural language understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D Manning, and Quoc V Le. 2019 · 1907
Earlier work this paper cites.
Self-knowledge distillation in natural language processing
Sangchul Hahn and Heeyoul Choi. 2019 · 1908
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang. 2019 · 1910
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Tutornet: Towards flexible knowledge distillation for end-to-end speech recognition
Ji Won Yoon, Hyeonseung Lee, Hyung Yong Kim, Won Ik Cho, and Nam Soo Kim. 2020 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. 2009 · 2009
Earlier work this paper cites.
Towards zero-shot knowledge distillation for natural language processing
Ahmad Rashid, Vasileios Lioutas, Abbas Ghaddar, and Mehdi Rezagholizadeh. 2020 · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik. 2015 · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Cited alongside, same era.
Distilling knowledge from ensembles of neural networks for speech recognition
Yevgen Chebotar and Austin Waters. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2017 · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. 2018 · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox. 2019 · 2019
Later among the works it cites.
Toward general scene graph: Integration of visual semantic knowledge with entity synset alignment
Woo Suk Choi, Kyoung-Woon On, Yu-Jung Heo, and Byoung-Tak Zhang. 2020 · 2020
Later among the works it cites.
Online knowledge distillation via collaborative learning
Qiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu, Ding Liang, Xiaolin Hu, and Ping Luo. 2020 · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak. 2020 · 2020
Later among the works it cites.
Why skip if you can combine: A simple knowledge distillation technique for intermediate layers
Yimeng Wu, Peyman Passban, Mehdi Rezagholizadeh, and Qun Liu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tommaso Furlanello, Zachary C Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018 · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina. 2019 · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu. 2019 · 2019
Cited alongside, same era.
Regularizing class-wise predictions via self-knowledge distillation
Sukmin Yun, Jongjin Park, Kimin Lee, and Jinwoo Shin. 2020 · 2020
Later among the works it cites.
Annealing knowledge distillation
Aref Jafari, Mehdi Rezagholizadeh, Pranav Sharma, and Ali Ghodsi. 2021 · 2021
Closest in time.
Active curriculum learning
Borna Jafarpour, Dawn Sepehr, and Nick Pogrebnyakov. 2021 · 2021
Closest in time.
Not far away, not so close: Sample efficient nearest neighbour data augmentation via MiniMax
Ehsan Kamalloo, Mehdi Rezagholizadeh, Peyman Passban, and Ali Ghodsi. 2021 · 2021
Closest in time.
ALP-KD: attention-based layer projection for knowledge distillation
Peyman Passban, Yimeng Wu, Mehdi Rezagholizadeh, and Qun Liu. 2021 · 2021
Closest in time.
MATE-KD: Masked adversarial TExt, a companion to knowledge distillation
Ahmad Rashid, Vasileios Lioutas, and Mehdi Rezagholizadeh. 2021 · 2021
Closest in time.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al. 2021 · 2021
Closest in time.