Fetching the paper…
Reading the bibliography…
We present Knowledge Distillation with Meta Learning (MetaDistil), a simple yet effective alternative to traditional knowledge distillation (KD) methods where the teacher model is fixed during training.
Stochastic estimation of the maximum of a regression function
Jack Kiefer, Jacob Wolfowitz, et al · 1952
Earlier work this paper cites.
A stochastic approximation method
Naresh K. Sinha and Michael P. Griscik · 1971
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Learner-centered teacher-student relationships are effective: A meta-analysis
Jeffrey Cornelius-White · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Student-centered learning in higher education
Gloria Brown Wright · 2011
Earlier work this paper cites.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez Colmenarejo, Matthew W. Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martínez-Rubio, Mark Schmidt, and Frank Wood · 2018
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, Seonguk Park, and Nojun Kwak · 2018
Cited alongside, same era.
Learning deep representations with probabilistic knowledge transfer
Nikolaos Passalis and Anastasios Tefas · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Cited alongside, same era.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas C. Damianou, Neil D. Lawrence, and Zhenwen Dai · 2019
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2019
Later among the works it cites.
Characterising bias in compressed models
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton · 2020
Later among the works it cites.
Metadistiller: Network self-boosting via meta-learned top-down distillation
Benlin Liu, Yongming Rao, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu · 2019
Cited alongside, same era.
DARTS: differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Cited alongside, same era.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Cited alongside, same era.
Haojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang, Yaliang Li, and Jun Huang · 2020
Later among the works it cites.
Hieu Pham, Zihang Dai, Qizhe Xie, Minh-Thang Luong, and Quoc V Le · 2020
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Bert-of-theseus: Compressing BERT by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou · 2020
Later among the works it cites.
BERT loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley, Ke Xu, and Furu Wei · 2020
Later among the works it cites.
Comparing kullback-leibler divergence and mean squared error loss in knowledge distillation
Taehyeon Kim, Jaehoon Oh, Nakyil Kim, Sangwook Cho, and Se-Young Yun · 2021
Closest in time.
Learning student-friendly teacher networks for knowledge distillation
Dae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim, and Bohyung Han · 2021
Closest in time.
Hieu Pham, Xinyi Wang, Yiming Yang, and Graham Neubig · 2021
Closest in time.
Learning from deep model via exploring local targets, 2021
Wenxian Shi, Yuxuan Song, Hao Zhou, Bohan Li, and Lei Li · 2021
Closest in time.
Improving sequence-to-sequence pre-training via sequence span rewriting
Wangchunshu Zhou, Tao Ge, Ke Xu, and Furu Wei · 2021
Closest in time.
A survey on model compression for natural language processing
Canwen Xu and Julian McAuley · 2022
Closest in time.