Fetching the paper…
Reading the bibliography…
Recent advances in pre-training huge models on large amounts of text through self supervision have obtained state-of-the-art results in various natural language processing tasks.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin · 1903
Earlier work this paper cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 1904
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 1906
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler · 2012
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian J. McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir D. Bourdev · 2014
Earlier work this paper cites.
Effective use of word order for text categorization with convolutional neural networks
Rie Johnson and Tong Zhang · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Cited alongside, same era.
Song Han, Huizi Mao, and William J. Dally · 2016
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M. Dai, and Ian J. Goodfellow · 2017
Later among the works it cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, and Quoc V. Le et al · 2016
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Closest in time.
Introducing distilbert, a distilled version of bert
Victor Sanh · 2019
Closest in time.
Patient knowledge distillation for bert model compression, 2019
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Closest in time.
Well-read students learn better: On the importance of pre-training compact models, 2019
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
XtremeDistil: Multi-stage distillation for massive multilingual models
Subhabrata Mukherjee and Ahmed Hassan Awadallah · 2020
Closest in time.