Fetching the paper…
Reading the bibliography…
Magnitude pruning is a widely used strategy for reducing model size in pure supervised learning; however, it is less effective in the transfer learning regime that has become standard for state-of-the-art natural language processing applications.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1989
Earlier work this paper cites.
Optimal brain surgeon: Extensions and performance comparisons
Babak Hassibi, David G. Stork, and Gregory J. Wolff · 1993
Earlier work this paper cites.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin · 2004
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Structured sparsity through convex optimization
Francis Bach, Rodolphe Jenatton, Julien Mairal, and Guillaume Obozinski · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
M. Horowitz · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir D. Bourdev · 2014
Earlier work this paper cites.
Proximal algorithms
Neal Parikh and Stephen P. Boyd · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William J. Dally · 2015
Earlier work this paper cites.
Dynamic network surgery for efficient dnns
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learning sparse neural networks through l0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs, 2017
Shankar Iyer, Nikhil Dandekar, and Kornel Csernai · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Piggyback: Adding multiple tasks to a single, fixed network by learning to mask
Arun Mallya and Svetlana Lazebnik · 2018
Cited alongside, same era.
Faster gaze prediction with dense networks and fisher pruning
Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszár · 2018
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2018
Cited alongside, same era.
Amc: Automl for model compression and acceleration on mobile devices
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han · 2018
Cited alongside, same era.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta · 2018
Transfer learning in natural language processing
Sebastian Ruder, Matthew E. Peters, Swabha Swayamdipta, and Thomas Wolf · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Reweighted proximal pruning for large-scale language representation
Fu-Ming Guo, Sijia Liu, Finlay S. Mungall, Xue Lian Lin, and Yanzhi Wang · 2019
Later among the works it cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Global sparse momentum sgd for pruning very deep neural networks
Xiaohan Ding, Guiguang Ding, Xiangxin Zhou, Yuchen Guo, Ji Liu, and Jungong Han · 2019
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Later among the works it cites.
Pruning a bert-based question answering model
J. Scott McCarley · 2019
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Y. Yin, Lifeng Shang, Xin Jiang, Xusong Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, G. Dziugaite, D. M. Roy, and Michael Carbin · 2020
Closest in time.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari · 2020
Closest in time.
Poor man’s bert: Smaller and faster transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov · 2020
Closest in time.
Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp
Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S. Morcos · 2020
Closest in time.
Train large, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez · 2020
Closest in time.
Compressing bert: Studying the effects of weight pruning on transfer learning
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
Causal mediation analysis for interpreting neural nlp: The case of gender bias
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber · 2020
Closest in time.