Fetching the paper…
Reading the bibliography…
Model compression by way of parameter pruning, quantization, or distillation has recently gained popularity as an approach for reducing the computational requirements of modern deep neural network models for NLP.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Efficient 8-bit quantization of transformer neural machine language translation model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla. 1989 · 1989
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Quantization
R.M. Gray and D.L. Neuhoff. 1998 · 1998
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Improving the speed of neural networks on cpus
Vincent Vanhoucke, Andrew Senior, and Mark Z. Mao. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J. Dally. 2015 · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromańska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina. 2017 · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017 · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, and Weinan E. 2017 · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. 2018 · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018 · 2018
Earlier work this paper cites.
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Cited alongside, same era.
Visualizing and understanding the effectiveness of bert
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
The algorithmic divide and equality in the age of artificial intelligence
Peter K. Yu. 2020 · 2020
Later among the works it cites.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Jiatao Gu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
The low-resource double bind: An empirical study of pruning for low-resource machine translation
Orevaoghene Ahia, Julia Kreutzer, and Sara Hooker. 2021 · 2021
Later among the works it cites.
BinaryBERT: Pushing the limit of BERT quantization
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. 2021 · 2021
Later among the works it cites.
EarlyBERT: Efficient BERT training via early-bird lottery tickets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019 · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Cited alongside, same era.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz. 2019 · 2019
Cited alongside, same era.
The de-democratization of ai: Deep learning and the compute divide in artificial intelligence research
Nur Ahmed and Muntasir Wahed. 2020 · 2020
Cited alongside, same era.
The generalization-stability tradeoff in neural network pruning
Brian Bartoldson, Ari Morcos, Adrian Barbu, and Gordon Erlebacher. 2020 · 2020
Cited alongside, same era.
Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu. 2021 · 2021
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee. 2021 · 2021
Later among the works it cites.
Multi-prize lottery ticket hypothesis: Finding accurate binary neural networks by pruning a randomly weighted network
James Diffenderfer and Bhavya Kailkhura. 2021 · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021 · 2021
Later among the works it cites.
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Parameter-efficient transfer learning with diff pruning
Demi Guo, Alexander Rush, and Yoon Kim. 2021 · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Robustness to pruning predicts generalization in deep neural networks
Lorenz Kuhn, Clare Lyle, Aidan N. Gomez, Jonas Rothfuss, and Yarin Gal. 2021 · 2021
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi. 2021 · 2021
Later among the works it cites.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush. 2021 · 2021
Later among the works it cites.
Super tickets in pre-trained language models: From model compression to improving generalization
Chen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
Lost in pruning: The effects of pruning neural networks beyond test accuracy
Lucas Liebenwein, Cenk Baykal, Brandon Carter, David Gifford, and Daniela Rus. 2021 · 2021
Later among the works it cites.
An empirical investigation of the role of pre-training in lifelong learning
Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, and Emma Strubell. 2021 · 2021
Later among the works it cites.
Beyond preserved accuracy: Evaluating loyalty and robustness of BERT compression
Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, and Furu Wei. 2021 · 2021
Later among the works it cites.
Sharpness-aware minimization improves language model generalization
Dara Bahri, Hossein Mobahi, and Yi Tay. 2022 · 2022
Closest in time.
Low-pass filtering sgd for recovering flat optima in the deep learning optimization landscape
Devansh Bisla, Jing Wang, and Anna Choromanska. 2022 · 2022
Closest in time.
Measuring the carbon intensity of ai in cloud instances
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan. 2022 · 2022
Closest in time.
Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask?
Mansheej Paul, Feng Chen, Brett W. Larsen, Jonathan Frankle, Surya Ganguli, and Gintare Karolina Dziugaite. 2022 · 2022
Closest in time.
Structured pruning learns compact and accurate models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Closest in time.