Fetching the paper…
Reading the bibliography…
Pre-trained universal feature extractors, such as BERT for natural language processing and VGG for computer vision, have become effective methods for improving deep learning models without requiring more labeled data.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Low-memory neural network training: A technical report
Nimit Sharad Sohoni, Christopher Richard Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré. 2019 · 1904
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
Luck matters: Understanding training dynamics of deep relu networks
Yuandong Tian, Tina Jiang, Qucheng Gong, and Ari S. Morcos. 2019 · 1905
Earlier work this paper cites.
KERMIT: generative insertion-based modeling for sequences
William Chan, Nikita Kitaev, Kelvin Guu, Mitchell Stern, and Jakob Uszkoreit. 2019 · 1906
Earlier work this paper cites.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan N. Gomez, and Erich Elsen. 2019 · 1906
Earlier work this paper cites.
Scalable syntax-aware language models using knowledge distillation
Adhiguna Kuncoro, Chris Dyer, Laura Rimell, Stephen Clark, and Phil Blunsom. 2019 · 1906
Earlier work this paper cites.
Ari S. Morcos, Haonan Yu, Michela Paganini, and Yuand ong Tian. 2019 · 1906
Earlier work this paper cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S. Morcos. 2019 · 1906
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke S. Zettlemoyer. 2019 · 1907
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Cited alongside, same era.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally. 2015 · 2015
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. 2018 · 2018
Later among the works it cites.
Reconciling modern machine learning practice and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mand al. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Learning sparse neural networks through l-0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Gregory S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian J. Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Józefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Gordon Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul A. Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda B. Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Cited alongside, same era.
Eie: Efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally. 2016 · 2016
Cited alongside, same era.
Opennmt: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush. 2017 · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta. 2017 · 2017
Cited alongside, same era.
The illustrated transformer
Jay Alammar. 2018 · 2018
Cited alongside, same era.
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Later among the works it cites.
Using pre-training can improve model robustness and uncertainty
Dan Hendrycks, Kimin Lee, and Mantas Mazeika. 2019 · 2019
Later among the works it cites.
Specializing word embeddings (for parsing) by information bottleneck
Xiang Lisa Li and Jason Eisner. 2019 · 2019
Later among the works it cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.