Fetching the paper…
Reading the bibliography…
The growing energy and performance costs of deep learning have driven the community to reduce the size of neural networks by selectively pruning components.
A Modular Benchmarking Infrastructure for High-Performance and Reproducible Deep Learning
Tal Ben-Nun, Maciej Besta, Simon Huber, Alexandros Nikolaos Ziogas, Daniel Peter, and Torsten Hoefler. 2019 · 1901
Earlier work this paper cites.
Error feedback fixes SignSGD and other gradient compression schemes. In Proceedings of the Thirty-sixth International Conference on Machine Learning
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U Stich, and Martin Jaggi. 2019 · 1901
Earlier work this paper cites.
A distributed synchronous SGD algorithm with global Top-k sparsification for low bandwidth networks. In 2019 IEEE 39th International Conference on Distributed Computing Systems Workshop on Networks
Shaohuai Shi, Qiang Wang, Kaiyong Zhao, Zhenheng Tang, Yuxin Wang, Xiang Huang, and Xiaowen Chu. 2019a · 1901
Earlier work this paper cites.
The State of Sparsity in Deep Neural Networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
Star-Transformer. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang. 2019b · 1902
Earlier work this paper cites.
Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
Hesham Mostafa and Xin Wang. 2019 · 1902
Earlier work this paper cites.
How Can We Be So Dense? The Benefits of Using Highly Sparse Representations
Subutai Ahmad and Luiz Scheinkman. 2019 · 1903
Earlier work this paper cites.
Stabilizing the Lottery Ticket Hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. 2020b · 1903
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations. In Proceedings of the Seventh International Conference on Learning Representations
Dan Hendrycks and Thomas Dietterich. 2019 · 1903
Earlier work this paper cites.
Communication-efficient distributed SGD with sketching. In Advances in Neural Information Processing Systems
Nikita Ivkin, Daniel Rothchild, Enayat Ullah, Ion Stoica, Raman Arora, et al · 1903
Earlier work this paper cites.
Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang, Stanley Osher, Yingyong Qi, and Jack Xin. 2019 · 1903
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure
Xiaohan Ding, Guiguang Ding, Yuchen Guo, and Jungong Han. 2019a · 1904
Earlier work this paper cites.
Learning Sparse Networks Using Targeted Dropout
Aidan N. Gomez, Ivan Zhang, Siddhartha Rao Kamalakara, Divyam Madaan, Kevin Swersky, Yarin Gal, and Geoffrey E. Hinton. 2019 · 1905
Earlier work this paper cites.
Sparse Transfer Learning via Winning Lottery Tickets
Rahul Mehta. 2019 · 1905
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan and Quoc V. Le. 2020 · 1905
Earlier work this paper cites.
DoubleSqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression. In Proceedings of the Thirty-sixth International Conference on Machine Learning
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu. 2019 · 1905
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 1905
Earlier work this paper cites.
Eigendamage: Structured pruning in the kronecker-factored eigenbasis
Chaoqi Wang, Roger Grosse, Sanja Fidler, and Guodong Zhang. 2019 · 1905
Earlier work this paper cites.
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2020 · 1905
Earlier work this paper cites.
The Generalization-Stability Tradeoff In Neural Network Pruning
Brian R. Bartoldson, Ari S. Morcos, Adrian Barbu, and Gordon Erlebacher. 2020 · 1906
Earlier work this paper cites.
Qsparse-local-SGD: Distributed SGD with quantization, sparsification, and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas N Diggavi. 2020 · 1906
Earlier work this paper cites.
A Signal Propagation Perspective for Pruning Neural Networks at Initialization
Namhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, and Philip H. S. Torr. 2020a · 1906
Earlier work this paper cites.
Importance Estimation for Neural Network Pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019 · 1906
Earlier work this paper cites.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Ari S. Morcos, Haonan Yu, Michela Paganini, and Yuandong Tian. 2019 · 1906
Earlier work this paper cites.
Discovering Neural Wirings
Mitchell Wortsman, Ali Farhadi, and Mohammad Rastegari. 2019 · 1906
Earlier work this paper cites.
Sparse Networks from Scratch: Faster Training without Losing Performance
Tim Dettmers and Luke Zettlemoyer. 2019 · 1907
Earlier work this paper cites.
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. 2019 · 1907
Earlier work this paper cites.
Large Memory Layers with Product Keys
Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2019 · 1907
Earlier work this paper cites.
Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
Yuanzhi Li, Colin Wei, and Tengyu Ma. 2020b · 1907
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Earlier work this paper cites.
MASR: A Modular Accelerator for Sparse RNNs
Udit Gupta, Brandon Reagen, Lillian Pentecost, Marco Donato, Thierry Tambe, Alexander M. Rush, Gu-Yeon Wei, and David Brooks. 2019 · 1908
Earlier work this paper cites.
Adversarial Neural Pruning with Latent Vulnerability Suppression
Divyam Madaan, Jinwoo Shin, and Sung Ju Hwang. 2020 · 1908
Earlier work this paper cites.
DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity Measures
Huanrui Yang, Wei Wen, and Hai Li. 2020b · 1908
Earlier work this paper cites.
Adaptively sparse transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
Gonçalo M Correia, Vlad Niculae, and André FT Martins. 2019 · 1909
Earlier work this paper cites.
Global Sparse Momentum SGD for Pruning Very Deep Neural Networks
Xiaohan Ding, Guiguang Ding, Xiangxin Zhou, Yuchen Guo, Jungong Han, and Ji Liu. 2019b · 1909
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout. In Proceedings of the Eighth International Conference on Learning Representations
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 1909
Earlier work this paper cites.
Reweighted proximal pruning for large-scale language representation
Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. 2019a · 1909
Earlier work this paper cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Richard G. Baraniuk, Zhangyang Wang, and Yingyan Lin. 2020 · 1909
Earlier work this paper cites.
Gate Decorator: Global Filter Pruning Method for Accelerating Deep Convolutional Neural Networks
Zhonghui You, Kun Yan, Jinmian Ye, Meng Ma, and Ping Wang. 2019 · 1909
Earlier work this paper cites.
MLPerf Training Benchmark
Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debojyoti Dutta, Udit Gupta, Kim Hazelwood, Andrew Hock, Xinyuan Huang, Atsushi Ike, Bill Jia, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Guokai Ma, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St. John, Tsuguchika Tabaru, Carole-Jean Wu, Lingjie Xu, Masafumi Yamazaki, Cliff Young, and Matei Zaharia. 2020 · 1910
Earlier work this paper cites.
Structured Pruning of a BERT-based Question Answering Model
J. S. McCarley, Rishav Chakravarti, and Avirup Sil. 2020 · 1910
Earlier work this paper cites.
SPEC2: SPECtral SParsE CNN Accelerator on FPGAs
Yue Niu, Hanqing Zeng, Ajitesh Srivastava, Kartik Lakhotia, Rajgopal Kannan, Yanzhi Wang, and Viktor Prasanna. 2019 · 1910
Earlier work this paper cites.
Structured pruning of large language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2020a · 1910
Earlier work this paper cites.
On the discrepancy between the theoretical analysis and practical implementations of compressed communication for distributed deep learning. In Proceedings of the AAAI Conference on Artificial Intelligence
Aritra Dutta, El Houcine Bergou, Ahmed M Abdelmoniem, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini, and Panos Kalnis. 2020 · 1911
Earlier work this paper cites.
Fast Sparse ConvNets
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan. 2019 · 1911
Earlier work this paper cites.
Rigging the Lottery: Making All Tickets Winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2020 · 1911
Earlier work this paper cites.
What Do Compressed Deep Neural Networks Forget?
Sara Hooker, Aaron Courville, Gregory Clark, Yann Dauphin, and Andrea Frome. 2019 · 1911
Earlier work this paper cites.
Provable Filter Pruning for Efficient Neural Networks
Lucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman, and Daniela Rus. 2020 · 1911
Earlier work this paper cites.
What’s Hidden in a Randomly Weighted Neural Network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari. 2020 · 1911
Earlier work this paper cites.
The Search for Sparse, Robust Neural Networks
Justin Cosentino, Federico Zaiter, Dan Pei, and Jun Zhu. 2019 · 1912
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. 2020a · 1912
Earlier work this paper cites.
Taxonomy and Evaluation of Structured Compression of Convolutional Neural Networks
Andrey Kuzmin, Markus Nagel, Saurabh Pitre, Sandeep Pendyam, Tijmen Blankevoort, and Max Welling. 2019 · 1912
Earlier work this paper cites.
Winning the Lottery with Continuous Sparsification
Pedro Savarese, Hugo Silva, and Michael Maire. 2020 · 1912
Earlier work this paper cites.
Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
Guangxiang Zhao, Junyang Lin, Zhiyuan Zhang, Xuancheng Ren, Qi Su, and Xu Sun. 2019 · 1912
Earlier work this paper cites.
The organization of behavior: A neuropsychological theory
Donald O. Hebb. 1949 · 1949
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014a · 1958
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014b · 1958
Earlier work this paper cites.
A Survey of Indexing Techniques for Sparse Matrices
Udo W. Pooch and Al Nieder. 1973 · 1973
Earlier work this paper cites.
Comparing Biases for Minimal Network Construction with Back-Propagation. In Advances in Neural Information Processing Systems
Stephen Hanson and Lorien Pratt. 1989 · 1988
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Michael C Mozer and Paul Smolensky. 1988 · 1988
Earlier work this paper cites.
Neural net pruning-why and how. In IEEE 1988 International Conference on Neural Networks
Sietsma and Dow. 1988 · 1988
Earlier work this paper cites.
A Back-Propagation Algorithm with Optimal Use of Hidden Units
Yves Chauvin. 1989 · 1989
Earlier work this paper cites.
Pruning versus clipping in neural networks
Steven A Janowsky. 1989 · 1989
Earlier work this paper cites.
A simple procedure for pruning back-propagation trained neural networks
E. D. Karnin. 1990 · 1990
Earlier work this paper cites.
Optimal Brain Damage
Yann Le Cun, John S. Denker, and Sara A. Solla. 1990 · 1990
Earlier work this paper cites.
The Evolution of Connectivity: Pruning Neural Networks Using Genetic Algorithms. In Proceedings of the International Joint Conference on Neural Networks
D. Whitley and C. Bogart. 1990 · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Automatic fusion and splitting of artificial neural elements in optimizing the network size. In Conference Proceedings 1991 IEEE International Conference on Systems, Man, and Cybernetics
K. Kameyama and Y. Kosugi. 1991 · 1991
Earlier work this paper cites.
A Simple Weight Decay Can Improve Generalization. In Proceedings of the 4th International Conference on Neural Information Processing Systems
Anders Krogh and John A. Hertz. 1991 · 1991
Earlier work this paper cites.
Note on generalization, regularization and architecture selection in nonlinear learning systems. In Neural Networks for Signal Processing Proceedings of the 1991 IEEE Workshop
John E Moody. 1991 · 1991
Earlier work this paper cites.
Creating artificial neural networks that generalize
Jocelyn Sietsma and Robert JF Dow. 1991 · 1991
Earlier work this paper cites.
Second Order Derivatives for Network Pruning: Optimal Brain Surgeon. In Advances in Neural Information Processing Systems 5, [NIPS Conference]
Babak Hassibi and David G. Stork. 1992 · 1992
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
Steven J Nowlan and Geoffrey E Hinton. 1992 · 1992
Earlier work this paper cites.
A pruning technique maximizing generalization. In Proceedings of 1993 International Conference on Neural Networks (IJCNN-93-Nagoya, Japan)
P. Burrascano. 1993 · 1993
Earlier work this paper cites.
Improving model selection by nonconvergent methods
William Finnoff, Ferdinand Hergert, and Hans Georg Zimmermann. 1993 · 1993
Earlier work this paper cites.
Removal of hidden units and weights for back propagation networks. In Proceedings of 1993 International Conference on Neural Networks (IJCNN-93-Nagoya, Japan)
Masafumi Hagiwara. 1993 · 1993
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights. In Proceedings of the sixth annual conference on Computational learning theory
Geoffrey E Hinton and Drew Van Camp. 1993 · 1993
Earlier work this paper cites.
Reduced-size neural networks through singular value decomposition and subset selection
P. P. Kanjilal, P. K. Dey, and D. N. Banerjee. 1993 · 1993
Earlier work this paper cites.
Pruning algorithms-a survey
R. Reed. 1993 · 1993
Earlier work this paper cites.
Determination of the number of redundant hidden units in a three-layered feedforward neural network. In Proceedings of 1993 International Conference on Neural Networks (IJCNN-93-Nagoya, Japan)
S. Tamura, M. Tateishi, M. Matumoto, and S. Akita. 1993 · 1993
Earlier work this paper cites.
GANNet: A Genetic Algorithm for Optimizing Topology and Weights in Neural Network Design. In Proceedings of the International Workshop on Artificial Neural Networks: New Trends in Neural Computation
David White and Panos A. Ligomenides. 1993 · 1993
Earlier work this paper cites.
Structural Adaptation and Generalization in Supervised Feed-Forward Networks
Joydeep Ghosh and Kagan Tumer. 1994 · 1994
Earlier work this paper cites.
A simple and effective method for removal of hidden units and weights
Masafumi Hagiwara. 1994 · 1994
Earlier work this paper cites.
Controlled growth of cascade correlation nets. In International Conference on Artificial Neural Networks
Lars Kai Hansen et al · 1994
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I Jordan and Robert A Jacobs. 1994 · 1994
Earlier work this paper cites.
Determining the significance of input parameters using sensitivity analysis. In International Workshop on Artificial Neural Networks
Andries Petrus Engelbrecht, Ian Cloete, and Jacek M Zurada. 1995 · 1995
Earlier work this paper cites.
Pruning with generalization based weight saliencies: lambda OBD, lambda OBS. In Advances in Neural Information Processing Systems
Morten Pedersen, Lars Hansen, and Jan Larsen. 1996 · 1995
Earlier work this paper cites.
Evaluating Pruning Methods. In National Chiao-Tung University
Georg Thimm and Emile Fiesler. 1995 · 1995
Earlier work this paper cites.
Bayesian Regularization and Pruning Using a Laplace Prior
P. M. Williams. 1995 · 1995
Earlier work this paper cites.
Variable selection with neural networks
Tautvydas Cibas, Françoise Fogelman Soulié, Patrick Gallinari, and Sarunas Raudys. 1996 · 1996
Earlier work this paper cites.
A sensitivity analysis algorithm for pruning feedforward neural networks. In Proceedings of International Conference on Neural Networks (ICNN’96)
A. P. Engelbrecht and I. Cloete. 1996 · 1996
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field. 1996 · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani. 1996 · 1996
Earlier work this paper cites.
An iterative pruning algorithm for feedforward neural networks
G. Castellano, A. M. Fanelli, and M. Pelillo. 1997 · 1997
Earlier work this paper cites.
Connection pruning with static and adaptive pruning schedules
Lutz Prechelt. 1997 · 1997
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks. In Fifth European Conference on Speech Communication and Technology
Nikko Ström. 1997 · 1997
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Shun-ichi Amari. 1998 · 1998
Earlier work this paper cites.
Optimizing the number of hidden nodes of a feedforward artificial neural network. In 1998 IEEE International Joint Conference on Neural Networks Proceedings. IEEE World Congress on Computational Intelligence (Cat. No.98CH36227)
L. Fletcher, V. Katkovnik, F. E. Steffens, and A. P. Engelbrecht. 1998 · 1998
Earlier work this paper cites.
Subset-based training and pruning of sigmoid neural networks
Guian Zhou and Jennie Si. 1999 · 1999
Earlier work this paper cites.
Variable selection using neural-network models
Giovanna Castellano and Anna Maria Fanelli. 2000 · 2000
Earlier work this paper cites.
Pruning of basis functions in nonlinear approximators
Hema Chandrasekaran, Hung-Han Chen, and Michael T. Manry. 2000 · 2000
Earlier work this paper cites.
A new pruning heuristic based on variance analysis of sensitivity information
A. P. Engelbrecht. 2001 · 2001
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Sparse Weight Activation Training
Md Aamir Raihan and Tor M. Aamodt. 2020 · 2001
Earlier work this paper cites.
Occam’s razor. In Advances in neural information processing systems
Carl Edward Rasmussen and Zoubin Ghahramani. 2001 · 2001
Earlier work this paper cites.
A simple neural network pruning algorithm with application to filter synthesis. In Neural Processing Letters
Kenji Suzuki, Isao Horiba, and Noboru Sugie. 2001 · 2001
Earlier work this paper cites.
Sparse Bayesian learning and the relevance vector machine
Michael E Tipping. 2001 · 2001
Earlier work this paper cites.
Distributed Variance Reduction with Optimal Communication
Peter Davies, Vijaykrishna Gurunathan, Niusha Moshrefi, Saleh Ashkboos, and Dan Alistarh. 2020 · 2002
Earlier work this paper cites.
The Early Phase of Neural Network Training
Jonathan Frankle, David J. Schwab, and Ari S. Morcos. 2020c · 2002
Earlier work this paper cites.
Compressing large-scale transformer-based models: A case study on BERT
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov. 2020 · 2002
Earlier work this paper cites.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning. In Proceedings of the 5th Workshop on Representation Learning for NLP
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Earlier work this paper cites.
Automatic Pruning for Quantized Neural Networks
Luis Guerra, Bohan Zhuang, Ian Reid, and Tom Drummond. 2020 · 2002
Earlier work this paper cites.
Pruning untrained neural networks: Principles and Analysis
Soufiane Hayou, Jean-Francois Ton, Arnaud Doucet, and Yee Whye Teh. 2020 · 2002
Earlier work this paper cites.
Soft Threshold Weight Reparameterization for Learnable Sparsity
Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi. 2020 · 2002
Earlier work this paper cites.
Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez. 2020a · 2002
Earlier work this paper cites.
Proving the Lottery Ticket Hypothesis: Pruning is All You Need
Eran Malach, Gilad Yehudai, Shai Shalev-Shwartz, and Ohad Shamir. 2020 · 2002
Earlier work this paper cites.
A primer in BERTology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2002
Earlier work this paper cites.
HYDRA: Pruning Adversarially Robust Neural Networks
Vikash Sehwag, Shiqi Wang, Prateek Mittal, and Suman Jana. 2020 · 2002
Earlier work this paper cites.
Top-k Training of GANs: Improving GAN Performance by Throwing Away Bad Samples. In Advances in Neural Information Processing Systems
Samarth Sinha, Zhengli Zhao, Anirudh Goyal, Colin A Raffel, and Augustus Odena. 2020 · 2002
Earlier work this paper cites.
SpArch: Efficient Architecture for Sparse Matrix Multiplication
Zhekai Zhang, Hanrui Wang, Song Han, and William J. Dally. 2020 · 2002
Earlier work this paper cites.
Learned Threshold Pruning
Kambiz Azarian, Yash Bhalgat, Jinwon Lee, and Tijmen Blankevoort. 2020 · 2003
Earlier work this paper cites.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020 · 2003
Earlier work this paper cites.
Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training
Namhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr, and Martin Jaggi. 2020b · 2003
Earlier work this paper cites.
SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, and Jiwei Li. 2020 · 2003
Earlier work this paper cites.
Comparing Rewinding and Fine-tuning in Neural Network Pruning
Alex Renda, Jonathan Frankle, and Michael Carbin. 2020 · 2003
Earlier work this paper cites.
Communication-efficient distributed deep learning: A comprehensive survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li. 2020 · 2003
Earlier work this paper cites.
Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection
Mao Ye, Chengyue Gong, Lizhen Nie, Denny Zhou, Adam Klivans, and Qiang Liu. 2020 · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Non-negative matrix factorization with sparseness constraints
Patrik O Hoyer. 2004 · 2004
Earlier work this paper cites.
WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
Sidak Pal Singh and Dan Alistarh. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners. In Advances in Neural Information Processing Systems
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005)
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Deep Learning for Post-Processing Ensemble Weather Forecasts
Peter Grönquist, Chengyuan Yao, Tal Ben-Nun, Nikoli Dryden, Peter Dueben, Shigang Li, and Torsten Hoefler. 2020 · 2005
Earlier work this paper cites.
Imaging input and output of neocortical networks in vivo
Jason N. D. Kerr, David Greenberg, and Fritjof Helmchen. 2005 · 2005
Earlier work this paper cites.
When BERT Plays the Lottery, All Tickets Are Winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2005
Earlier work this paper cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020 · 2005
Earlier work this paper cites.
Bayesian Bits: Unifying Quantization and Pruning
Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling. 2020 · 2005
Earlier work this paper cites.
Directional Pruning of Deep Neural Networks
Shih-Kang Chao, Zhanyu Wang, Yue Xing, and Guang Cheng. 2020 · 2006
Cited alongside, same era.
High performance convolutional neural networks for document processing
Kumar Chellapilla, Sidd Puri, and Patrice Simard. 2006 · 2006
Cited alongside, same era.
ESPN: Extremely Sparse Pruned Networks
Minsu Cho, Ameya Joshi, and Chinmay Hegde. 2020 · 2006
Cited alongside, same era.
Progressive Skeletonization: Trimming more fat from a network at initialization
Pau de Jorge, Amartya Sanyal, Harkirat S. Behl, Philip H. S. Torr, Gregory Rogez, and Puneet K. Dokania. 2020 · 2006
Cited alongside, same era.
Sparse GPU Kernels for Deep Learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020 · 2006
Cited alongside, same era.
Cognitive and neural plasticity in aging: general and task-specific limitations
Sari Jones, Lars Nyberg, Johan Sandblom, Anna Stigsdotter Neely, Martin Ingvar, Karl Magnus Petersson, and Lars Bäckman. 2006 · 2006
Learning Efficient Convolutional Networks through Network Slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017 · 2017
Later among the works it cites.
Bayesian Compression for Deep Learning
Christos Louizos, Karen Ullrich, and Max Welling. 2017 · 2017
Later among the works it cites.
ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin. 2017 · 2017
Later among the works it cites.
A Tutorial on Fisher Information
Alexander Ly, Maarten Marsman, Josine Verhagen, Raoul Grasman, and Eric-Jan Wagenmakers. 2017 · 2017
Later among the works it cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Later among the works it cites.
Exploring the Regularity of Sparse Structure in Convolutional Neural Networks
Huizi Mao, Song Han, Jeff Pool, Wenshuo Li, Xingyu Liu, Yu Wang, and William J. Dally. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A node pruning algorithm based on a Fourier amplitude sensitivity test method
Philippe Lauret, Eric Fock, and Thierry Alex Mara. 2006 · 2006
Cited alongside, same era.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2006
Cited alongside, same era.
Dynamic Model Pruning with Feedback
Tao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev, and Martin Jaggi. 2020 · 2006
Cited alongside, same era.
Finding trainable sparse networks through Neural Tangent Transfer
Tianlin Liu and Friedemann Zenke. 2020 · 2006
Cited alongside, same era.
Predictive Coding Approximates Backprop along Arbitrary Computation Graphs
Beren Millidge, Alexander Tschantz, and Christopher L. Buckley. 2020 · 2006
Cited alongside, same era.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin. 2006 · 2006
Cited alongside, same era.
Later among the works it cites.
Diversity Networks: Neural Network Compression Using Determinantal Point Processes
Zelda Mariet and Suvrit Sra. 2017 · 2017
Later among the works it cites.
WRPN: Wide Reduced-Precision Networks
Asit K. Mishra, Eriko Nurvitadhi, Jeffrey J. Cook, and Debbie Marr. 2017 · 2017
Later among the works it cites.
Variational Dropout Sparsifies Deep Neural Networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Later among the works it cites.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2017 · 2017
Later among the works it cites.
Exploring Sparsity in Recurrent Neural Networks
Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta. 2017 · 2017
Later among the works it cites.
Structured Bayesian Pruning via Log-Normal Multiplicative Noise
Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Later among the works it cites.
A regularized framework for sparse and structured neural attention. In Advances in neural information processing systems
Vlad Niculae and Mathieu Blondel. 2017 · 2017
Later among the works it cites.
SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. 2017 · 2017
Later among the works it cites.
Faster CNNs with Direct Sparse Convolutions and Guided Pruning
Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey. 2017 · 2017
Later among the works it cites.
Routing Networks: Adaptive Selection of Non-linear Functions for Multi-Task Learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer. 2017 · 2017
Later among the works it cites.
Group sparse regularization for deep neural networks
Simone Scardapane, Danilo Comminiello, Amir Hussain, and Aurelio Uncini. 2017 · 2017
Later among the works it cites.
The Incredible Shrinking Neural Network: New Perspectives on Learning Representations Through The Lens of Pruning
Aditya Sharma, Nikolas Wolfe, and Bhiksha Raj. 2017 · 2017
Later among the works it cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Later among the works it cites.
Opening the Black Box of Deep Neural Networks via Information
Ravid Shwartz-Ziv and Naftali Tishby. 2017 · 2017
Later among the works it cites.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting. In Proceedings of the Thirty-Fourth International Conference on Machine Learning
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang. 2017 · 2017
Later among the works it cites.
Distributed mean estimation with limited communication. In International Conference on Machine Learning
Ananda Theertha Suresh, X Yu Felix, Sanjiv Kumar, and H Brendan McMahan. 2017 · 2017
Later among the works it cites.
Efficient Processing of Deep Neural Networks: A Tutorial and Survey
V. Sze, Y. Chen, T. Yang, and J. S. Emer. 2017 · 2017
Later among the works it cites.
Soft Weight-Sharing for Neural Network Compression
Karen Ullrich, Edward Meeds, and Max Welling. 2017 · 2017
Later among the works it cites.
Trends in Data Locality Abstractions for HPC Systems
Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fuerlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H J Kelly, Vitus Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, , and Miquel Pericas. 2017 · 2017
Later among the works it cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Minimal Effort Back Propagation for Convolutional Neural Networks
Bingzhen Wei, Xu Sun, Xuancheng Ren, and Jingjing Xu. 2017 · 2017
Later among the works it cites.
Second-order Optimization for Deep Reinforcement Learning using Kronecker-factored Approximation. In NIPS
Yuhuai Wu, Elman Mansimov, Roger B. Grosse, Shun Liao, and Jimmy Ba. 2017 · 2017
Later among the works it cites.
Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning
Tien-Ju Yang, Yu-Hsin Chen, and Vivienne Sze. 2017 · 2017
Later among the works it cites.
Scalpel: Customizing dnn pruning to the underlying hardware parallelism
Jiecao Yu, Andrew Lukefahr, David Palframan, Ganesh Dasika, Reetuparna Das, and Scott Mahlke. 2017 · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2017 · 2017
Later among the works it cites.
Learning Efficient Tensor Representations with Ring Structure Networks
Qibin Zhao, Masashi Sugiyama, and Andrzej Cichocki. 2017 · 2017
Later among the works it cites.
SparseNN: An Energy-Efficient Neural Network Accelerator Exploiting Input and Output Sparsity
Jingyang Zhu, Jingbo Jiang, Xizi Chen, and Chi-Ying Tsui. 2017 · 2017
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta. 2017 · 2017
Later among the works it cites.
The convergence of sparsified gradient methods. In Advances in Neural Information Processing Systems
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli. 2018 · 2018
Later among the works it cites.
Data-dependent coresets for compressing neural networks with applications to generalization bounds
Cenk Baykal, Lucas Liebenwein, Igor Gilitschenski, Dan Feldman, and Daniela Rus. 2018 · 2018
Later among the works it cites.
Deep Rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein. 2018 · 2018
Later among the works it cites.
Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
Tal Ben-Nun and Torsten Hoefler. 2018 · 2018
Later among the works it cites.
Benchmark Analysis of Representative Deep Neural Network Architectures
Simone Bianco, Remi Cadene, Luigi Celona, and Paolo Napoletano. 2018 · 2018
Later among the works it cites.
"Learning-Compression" Algorithms for Neural Net Pruning. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
M. A. Carreira-Perpinan and Y. Idelbayev. 2018 · 2018
Later among the works it cites.
Compressing Neural Networks using the Variational Information Bottleneck
Bin Dai, Chen Zhu, and David Wipf. 2018b · 2018
Later among the works it cites.
NeST: A Neural Network Synthesis Tool Based on a Grow-and-Prune Paradigm
Xiaoliang Dai, Hongxu Yin, and Niraj K. Jha. 2018a · 2018
Later among the works it cites.
DropBlock: A regularization method for convolutional networks. In Advances in Neural Information Processing Systems
Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. 2018 · 2018
Later among the works it cites.
Combating Adversarial Attacks Using Sparse Representations
Soorya Gopalakrishnan, Zhinus Marzi, Upamanyu Madhow, and Ramtin Pedarsani. 2018 · 2018
Later among the works it cites.
Morphnet: Fast & simple resource-constrained structure learning of deep networks. In Proceedings of the IEEE conference on computer vision and pattern recognition
Ariel Gordon, Elad Eban, Ofir Nachum, Bo Chen, Hao Wu, Tien-Ju Yang, and Edward Choi. 2018 · 2018
Later among the works it cites.
DNN Feature Map Compression using Learned Representation over GF (2). In Proceedings of the European Conference on Computer Vision (ECCV)
Denis Gudovskiy, Alec Hodgkinson, and Luca Rigazio. 2018 · 2018
Later among the works it cites.
Sparse dnns with improved adversarial robustness. In Advances in neural information processing systems
Yiwen Guo, Chao Zhang, Changshui Zhang, and Yurong Chen. 2018 · 2018
Later among the works it cites.
Data-Driven Sparse Structure Selection for Deep Neural Networks
Zehao Huang and Naiyan Wang. 2018 · 2018
Later among the works it cites.
A linear speedup analysis of distributed deep learning with sparse and quantized communication. In Advances in Neural Information Processing Systems
Peng Jiang and Gagan Agrawal. 2018 · 2018
Later among the works it cites.
Randomized distributed mean estimation: Accuracy vs. communication
Jakub Konečnỳ and Peter Richtárik. 2018 · 2018
Later among the works it cites.
Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
H. T. Kung, Bradley McDanel, and Sai Qian Zhang. 2018 · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training. In Proceedings of the Sixth International Conference on Learning Representations
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally. 2018 · 2018
Later among the works it cites.
Open subtitles 2018: Statistical rescoring of sentence alignments in large, noisy parallel corpora. In LREC 2018, Eleventh International Conference on Language Resources and Evaluation
Pierre Lison, Jörg Tiedemann, Milen Kouylekov, et al · 2018
Later among the works it cites.
Dynamic Deep Neural Networks: Optimizing Accuracy-Efficiency Trade-offs by Selective Execution
Lanlan Liu and Jia Deng. 2018 · 2018
Later among the works it cites.
Bayesian sparsification of gated recurrent neural networks
Ekaterina Lobacheva, Nadezhda Chirkova, and Dmitry Vetrov. 2018 · 2018
Later among the works it cites.
Learning Sparse Neural Networks through L 0 L_{0} Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
Sparse and constrained attention for neural machine translation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
Chaitanya Malaviya, Pedro Ferreira, and André FT Martins. 2018 · 2018
Later among the works it cites.
Automated Pruning for Deep Neural Network Compression
Franco Manessi, Alessandro Rozza, Simone Bianco, Paolo Napoletano, and Raimondo Schettini. 2018 · 2018
Later among the works it cites.
An Empirical Model of Large-Batch Training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. 2018 · 2018
Later among the works it cites.
Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural Networks
Deepak Mittal, Shweta Bhardwaj, Mitesh M. Khapra, and Balaraman Ravindran. 2018 · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta. 2018 · 2018
Later among the works it cites.
Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro. 2018 · 2018
Later among the works it cites.
Image Transformer. In International Conference on Machine Learning
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Later among the works it cites.
Deep Expander Networks: Efficient Deep Networks from Graph Theory
Ameya Prabhu, Girish Varma, and Anoop Namboodiri. 2018 · 2018
Later among the works it cites.
Compressing DMA engine: Leveraging activation sparsity for training deep neural networks. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)
Minsoo Rhu, Mike O’Connor, Niladrish Chatterjee, Jeff Pool, Youngeun Kwon, and Stephen W Keckler. 2018 · 2018
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le. 2018 · 2018
Later among the works it cites.
Sparsified SGD with memory. In Advances in Neural Information Processing Systems
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi. 2018 · 2018
Later among the works it cites.
Learning Sparse Neural Networks via Sensitivity-Driven Regularization
Enzo Tartaglione, Skjalg Lepsøy, Attilio Fiandrotti, and Gianluca Francini. 2018 · 2018
Later among the works it cites.
Faster gaze prediction with dense networks and Fisher pruning
Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszár. 2018 · 2018
Later among the works it cites.
Variance-based gradient compression for efficient distributed deep learning. In Proceedings of the Sixth International Conference on Learning Representations, Workshop Track
Yusuke Tsuzuku, Hiroto Imachi, and Takuya Akiba. 2018 · 2018
Later among the works it cites.
ATOMO: Communication-efficient learning via atomic sparsification. In Advances in Neural Information Processing Systems
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright. 2018 · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization. In Advances in Neural Information Processing Systems
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Rethinking the smaller-norm-less-informative assumption in channel pruning of convolution layers
Jianbo Ye, Xin Lu, Zhe Lin, and James Z Wang. 2018 · 2018
Later among the works it cites.
NISP: Pruning Networks using Neuron Importance Score Propagation
Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry S. Davis. 2018 · 2018
Later among the works it cites.
Learning strict identity mappings in deep residual networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Xin Yu, Zhiding Yu, and Srikumar Ramalingam. 2018 · 2018
Later among the works it cites.
Cambricon-S: Addressing Irregularity in Sparse Neural Networks through A Cooperative Software/Hardware Approach. In 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen. 2018 · 2018
Later among the works it cites.
Critical Learning Periods in Deep Neural Networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto. 2019 · 2019
Later among the works it cites.
A Convergence Theory for Deep Learning via Over-Parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. 2019 · 2019
Later among the works it cites.
Eyeriss v2: A Flexible Accelerator for Emerging Deep Neural Networks on Mobile Devices
Yu-Hsin Chen, Tien-Ju Yang, Joel Emer, and Vivienne Sze. 2019 · 2019
Later among the works it cites.
Fine-tune BERT with Sparse Self-Attention Mechanism. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
Baiyun Cui, Yingming Li, Ming Chen, and Zhongfei Zhang. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Pre-Defined Sparse Neural Networks With Hardware Acceleration
S. Dey, K. Huang, P. A. Beerel, and K. M. Chugg. 2019 · 2019
Later among the works it cites.
Exploiting the input sparsity to accelerate deep neural networks: poster. In Proceedings of the 24th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, PPoPP 2019, Washington, DC, USA, February 16-20, 2019
Xiao Dong, Lei Liu, Guangli Li, Jiansong Li, Peng Zhao, Xueying Wang, and Xiaobing Feng. 2019 · 2019
Later among the works it cites.
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. 2019 · 2019
Later among the works it cites.
Neural Architecture Search: A Survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019 · 2019
Later among the works it cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Later among the works it cites.
Accelerating Convolutional Neural Networks via Activation Map Compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Georgios Georgiadis. 2019 · 2019
Later among the works it cites.
Full deep neural network training on a pruned weight budget
Maximilian Golub, Guy Lemieux, and Mieszko Lis. 2019 · 2019
Later among the works it cites.
SparTen: A Sparse Tensor Accelerator for Convolutional Neural Networks. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture
Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and T. N. Vijaykumar. 2019 · 2019
Later among the works it cites.
AMC: AutoML for Model Compression and Acceleration on Mobile Devices
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. 2019a · 2019
Later among the works it cites.
Filter Pruning via Geometric Median for Deep Convolutional Neural Networks Acceleration
Yang He, Ping Liu, Ziwei Wang, Zhilan Hu, and Yi Yang. 2019b · 2019
Later among the works it cites.
ExTensor: An Accelerator for Sparse Tensor Algebra. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture
Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W. Fletcher. 2019 · 2019
Later among the works it cites.
Optimal Sparsity-Sensitive Bounds for Distributed Mean Estimation. In Advances in Neural Information Processing Systems
Ziyue Huang, Wang Yilei, Ke Yi, et al · 2019
Later among the works it cites.
The IWSLT 2019 evaluation campaign. In 16th International Workshop on Spoken Language Translation 2019
Niehues Jan, Roldano Cattoni, Stuker Sebastian, Matteo Negri, Marco Turchi, Salesky Elizabeth, Sanabria Ramon, Barrault Loic, Specia Lucia, and Marcello Federico. 2019 · 2019
Later among the works it cites.
DeepSZ: A Novel Framework to Compress Deep Neural Networks by Using Error-Bounded Lossy Compression. In Proceedings of the 28th International Symposium on High-Performance Parallel and Distributed Computing
Sian Jin, Sheng Di, Xin Liang, Jiannan Tian, Dingwen Tao, and Franck Cappello. 2019 · 2019
Later among the works it cites.
Efficient Language Modeling with Automatic Relevance Determination in Recurrent Neural Networks. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)
Maxim Kodryan, Artem Grachev, Dmitry Ignatov, and Dmitry Vetrov. 2019 · 2019
Later among the works it cites.
Limitations of the empirical Fisher approximation for natural gradient descent. In Advances in Neural Information Processing Systems
Frederik Kunstner, Philipp Hennig, and Lukas Balles. 2019 · 2019
Later among the works it cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Later among the works it cites.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr. 2019 · 2019
Later among the works it cites.
SqueezeFlow: A Sparse CNN Accelerator Exploiting Concise Convolution Rules
J. Li, S. Jiang, S. Gong, J. Wu, J. Yan, G. Yan, and X. Li. 2019 · 2019
Later among the works it cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2019 · 2019
Later among the works it cites.
3LC: Lightweight and Effective Traffic Compression for Distributed Machine Learning. In Proceedings of the Conference on Systems and Machine Learning
Hyeontaek Lim, David Andersen, and Michael Kaminsky. 2019 · 2019
Later among the works it cites.
Dynamic Sparse Graph for Efficient Deep Learning
Liu Liu, Lei Deng, Xing Hu, Maohua Zhu, Guoqi Li, Yufei Ding, and Yuan Xie. 2019 · 2019
Later among the works it cites.
Rethinking the Value of Network Pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2019b · 2019
Later among the works it cites.
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Later among the works it cites.
AutoPruner: An End-to-End Trainable Filter Pruning Method for Efficient Deep Model Inference
Jian-Hao Luo and Jianxin Wu. 2019 · 2019
Later among the works it cites.
PruneTrain
Sangkug Lym, Esha Choukse, Siavash Zangeneh, Wei Wen, Sujay Sanghavi, and Mattan Erez. 2019 · 2019
Later among the works it cites.
Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka. 2019 · 2019
Later among the works it cites.
SparCML: High-performance sparse communication for machine learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Cèdric Renggli, Saleh Ashkboos, Mehdi Aghagolzadeh, Dan Alistarh, and Torsten Hoefler. 2019 · 2019
Later among the works it cites.
Filter Distillation for Network Compression
Xavier Suau, Luca Zappella, and Nicholas Apostoloff. 2019 · 2019
Later among the works it cites.
Sparse gradient compression for distributed SGD. In International Conference on Database Systems for Advanced Applications
Haobo Sun, Yingxia Shao, Jiawei Jiang, Bin Cui, Kai Lei, Yu Xu, and Jiang Wang. 2019 · 2019
Later among the works it cites.
MnasNet: Platform-Aware Neural Architecture Search for Mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the Seventh International Conference on Learning Representations
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019 · 2019
Later among the works it cites.
AutoPrune: Automatic Network Pruning by Regularizing Auxiliary Parameters. In Advances in Neural Information Processing Systems
Xia Xiao, Zigeng Wang, and Sanguthevar Rajasekaran. 2019 · 2019
Later among the works it cites.
MLPrune: Multi-Layer Pruning for Automated Neural Network Compression
Wenyuan Zeng and Raquel Urtasun. 2019 · 2019
Later among the works it cites.
Eager Pruning: Algorithm and Architecture Support for Fast Training of Deep Neural Networks. In Proceedings of the 46th International Symposium on Computer Architecture
Jiaqi Zhang, Xiangru Chen, Mingcong Song, and Tao Li. 2019 · 2019
Later among the works it cites.
SNAP: A 1.67 21.55TOPS/W Sparse Neural Acceleration Processor for Unstructured Sparse Deep Neural Network Inference in 16nm CMOS
Jie-Fang Zhang, Ching-En Lee, C. Liu, Y. Shao, Stephen W. Keckler, and Zhengya Zhang. 2019a · 2019
Later among the works it cites.
CompAct: On-Chip ComPression of ActIvations for Low Power Systolic Array Based CNN Acceleration
Jeff (Jun) Zhang, Parul Raj, Shuayb Zarar, Amol Ambardekar, and Siddharth Garg. 2019b · 2019
Later among the works it cites.
Discrimination-aware Channel Pruning for Deep Neural Networks
Zhuangwei Zhuang, Mingkui Tan, Bohan Zhuang, Jing Liu, Yong Guo, Qingyao Wu, Junzhou Huang, and Jinhui Zhu. 2019 · 2019
Later among the works it cites.
Interval Adjoint Significance Analysis for Neural Networks. In International Conference on Computational Science
Sher Afghan and Uwe Naumann. 2020 · 2020
Later among the works it cites.
Storage Efficient and Dynamic Flexible Runtime Channel Pruning via Deep Reinforcement Learning
Jianda Chen, Shangyu Chen, and Sinno Jialin Pan. 2020 · 2020
Later among the works it cites.
A Survey of Model Compression and Acceleration for Deep Neural Networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2020 · 2020
Later among the works it cites.
A comprehensive survey on model compression and acceleration
Tejalal Choudhary, Vipul Mishra, Anurag Goswami, and Jagannathan Sarangapani. 2020 · 2020
Later among the works it cites.
Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
L. Deng, G. Li, S. Han, L. Shi, and Y. Xie. 2020 · 2020
Later among the works it cites.
Top-KAST: Top-K Always Sparse Training
Siddhant Jayakumar, Razvan Pascanu, Jack Rae, Simon Osindero, and Erich Elsen. 2020 · 2020
Later among the works it cites.
Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning
Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. 2020 · 2020
Later among the works it cites.
Backpropagation and the brain
Timothy P Lillicrap, Adam Santoro, Luke Marris, Colin J Akerman, and Geoffrey Hinton. 2020 · 2020
Later among the works it cites.
Reuse Kernels or Activations? A Flexible Dataflow for Low-Latency Spectral CNN Acceleration. In Proceedings of the 2020 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
Yue Niu, Rajgopal Kannan, Ajitesh Srivastava, and Viktor Prasanna. 2020 · 2020
Later among the works it cites.
NVIDIA A100 Tensor Core GPU Architecture
Nvidia. 2020 · 2020
Later among the works it cites.
SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna. 2020 · 2020
Later among the works it cites.
Robust Sparse Regularization: Defending Adversarial Attacks Via Regularized Sparse Network. In Proceedings of the 2020 on Great Lakes Symposium on VLSI
Adnan Siraj Rakin, Zhezhi He, Li Yang, Yanzhi Wang, Liqiang Wang, and Deliang Fan. 2020 · 2020
Later among the works it cites.
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig. 2020 · 2020
Later among the works it cites.
DropNet: Reducing Neural Network Complexity via Iterative Pruning. In Proceedings of the 37th International Conference on Machine Learning
Chong Min John Tan and Mehul Motani. 2020 · 2020
Later among the works it cites.
O ( n ) O(n) Connections are Expressive Enough: Universal Approximability of Sparse Transformers. In Advances in Neural Information Processing Systems
Chulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar. 2020 · 2020
Later among the works it cites.
Neuron-level Structured Pruning using Polarization Regularizer
Tao Zhuang, Zhixuan Zhang, Yuheng Huang, Xiaoyi Zeng, Kai Shuang, and Xiang Li. 2020 · 2020
Later among the works it cites.
Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
Less bits is more: How pruning deep binary networks increases weight capacity
Yunqiang Li, Silvia Laura Pintea, and Jan van Gemert. 2021 · 2021
Closest in time.