Fetching the paper…
Reading the bibliography…
While larger neural models are pushing the boundaries of what deep learning can do, often more weights are needed to train models rather than to run inference for tasks.
Thirteen ways to look at the correlation coefficient
Joseph Lee Rodgers and W. Alan Nicewander · 1988
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G. Stork · 1993
Earlier work this paper cites.
Pruning algorithms-a survey
R. Reed · 1993
Earlier work this paper cites.
An iterative pruning algorithm for feedforward neural networks
Giovanna Castellano, Anna Maria Fanelli, and Marcello Pelillo · 1997
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gerard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
SSD: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C. Berg · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Block-sparse recurrent neural networks
Sharan Narang, Eric Undersander, and Gregory Diamos · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin · 2017
Earlier work this paper cites.
Faster cnns with direct sparse convolutions and guided pruning
Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey · 2017
Earlier work this paper cites.
GPU kernels for block-sparse weights
S. Gray, A. Radford, and D. P. Kingma · 2017
Earlier work this paper cites.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2017
Earlier work this paper cites.
Dilated residual networks
Fisher Yu, Vladlen Koltun, and Thomas Funkhouser · 2017
Earlier work this paper cites.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2018
Cited alongside, same era.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Cited alongside, same era.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
M. Zhu and S. Gupta · 2018
Cited alongside, same era.
Stacked U-Nets: A no-frills approach to natural image segmentation
Sohil Shah, Pallabi Ghosh, Larry S. Davis, and Tom Goldstein · 2018
Cited alongside, same era.
MobileNetV2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B. Brown, Prafulla Dhariwal, Scott Gray, Chris Hallacy, Benjamin Mann, Alec Radford, Aditya Ramesh, Nick Ryder, Daniel M. Ziegler, John Schulman, Dario Amodei, and Sam McCandlish · 2020
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
The computational limits of deep learning
Neil C. Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F. Manso · 2020
Later among the works it cites.
Train big, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joey Gonzalez · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-resolution image synthesis and semantic manipulation with conditional gans
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Cited alongside, same era.
Video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A Radford, K Narasimhan, T Salimans, and I Sutskever · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
SNIP: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Comparing rewinding and fine-tuning in neural network pruning
Alex Renda, Jonathan Frankle, and Michael Carbin · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, and Surya Ganguli · 2020
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Later among the works it cites.
Top-KAST: Top-k always sparse training
Siddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero, and Erich Elsen · 2020
Later among the works it cites.
Sparse weight activation training
Md Aamir Raihan and Tor M. Aamodt · 2020
Later among the works it cites.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2020
Later among the works it cites.
Dynamic sparse training: Find efficient sparse network from scratch with trainable masked layers
Junjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung, and Hayden K.H. So · 2020
Later among the works it cites.
Sparse gpu kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Later among the works it cites.
Winning the lottery with continuous sparsification
Pedro Savarese, Hugo Silva, and Michael Maire · 2020
Later among the works it cites.
Gradient flow in sparse neural networks and how lottery tickets win
Utku Evci, Yani A. Ioannou, Cem Keskin, and Yann Dauphin · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Imaginaire
NVIDIA · 2020
Later among the works it cites.
Deep learning examples
NVIDIA · 2020
Later among the works it cites.
Megatron-LM
NVIDIA · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Closest in time.
Pruning neural networks at initialization: Why are we missing the mark?
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2021
Closest in time.
Accelerating sparse deep neural networks
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Closest in time.
Learning n:m fine-grained structured sparse neural networks from scratch
A. Zhou, Y. Ma, J. Zhu, J. Liu, Z. Zhang, K. Yuan, W. Sun, , and H. Li · 2021
Closest in time.
Accelerated sparse neural training: A provable and efficient method to find n:m transposable masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Seffi Naor, and Daniel Soudry · 2021
Closest in time.
Do we actually need dense over-parameterization? In-time over-parameterization in sparse training
Shiwei Liu, Lu Yin, Decebal Constantin Mocanu, and Mykola Pechenizkiy · 2021
Closest in time.