Fetching the paper…
Reading the bibliography…
While pre-training and fine-tuning, e.g., BERT~\citep{devlin2018bert}, GPT-2~\citep{radford2019language}, have achieved great success in language understanding and generation tasks, the pre-trained models are usually too big for online deployment in terms of both memory cost and inference speed, which hinders them from practical online usage.
Model compression with multi-task knowledge distillation for web-scale question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang · 1904
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le · 1906
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Kl-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition
Dong Yu, Kaisheng Yao, Hang Su, Gang Li, and Frank Seide · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
First quora dataset release: Question pairs
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Xu Lan, Xiatian Zhu, and Shaogang Gong · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Mass: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Knowledge distillation in generations: More tolerant teachers educate better students
Chenglin Yang, Lingxi Xie, Siyuan Qiao, and Alan Yuille · 2018
Cited alongside, same era.
An efficient way to learn rules for grapheme-to-phoneme conversion in chinese
Zi-Rong Zhang, Min Chu, and Eric Chang · 2018
Cited alongside, same era.
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Later among the works it cites.
Multilingual neural machine translation with knowledge distillation
Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu · 2019
Later among the works it cites.
Distilling task-specific knowledge from bert into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin · 2019
Later among the works it cites.