Fetching the paper…
Reading the bibliography…
Deep and large pre-trained language models are the state-of-the-art for various natural language processing tasks.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019 · 1904
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Ali Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian J. McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Parsing with compositional vector grammars
Richard Socher, John Bauer, Christopher D. Manning, and Andrew Y. Ng. 2013 · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir D. Bourdev. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
Song Han, Huizi Mao, and William J. Dally. 2016 · 2016
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Training compact models for low resource entity tagging using pre-trained language models
Peter Izsak, Shira Guskin, and Moshe Wasserblat. 2019 · 2019
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
Effective dimensionality reduction for word embeddings
Vikas Raunak, Vivek Gupta, and Florian Metze. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Knowledge distillation from internal representations
Gustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Yao, Xing Fan, and Edward Guo. 2019 · 2019
Cited alongside, same era.
Bam! born-again multi-task networks for natural language understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D. Manning, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Introducing distilbert, a distilled version of bert
Victor Sanh. 2019 · 2019
Later among the works it cites.
Knowledge distillation for recurrent neural network language modeling with trust regularization
Yangyang Shi, Mei-Yuh Hwang, Xin Lei, and Haoyu Sheng. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Small and practical bert models for sequence labeling
Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Xin Li, and Amelia Archer. 2019 · 2019
Later among the works it cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Extreme language model compression with optimal subwords and shared projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 2019
Later among the works it cites.
PANLP at MEDIQA 2019: Pre-trained language models, transfer learning and knowledge distillation
Wei Zhu, Xiaofeng Zhou, Keqiang Wang, Xun Luo, Xiepeng Li, Yuan Ni, and Guotong Xie. 2019 · 2019
Later among the works it cites.