Fetching the paper…
Reading the bibliography…
Sparse neural networks are becoming increasingly important as the field seeks to improve the performance of existing models by scaling them up, while simultaneously trying to reduce power consumption and computational footprint.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 1902
Earlier work this paper cites.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan N. Gomez, and Erich Elsen · 1906
Earlier work this paper cites.
Non-differentiable supervised learning with evolution strategies and hybrid methods
Karel Lenc, Erich Elsen, Tom Schaul, and Karen Simonyan · 1906
Earlier work this paper cites.
Optimal Brain Damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Evaluating pruning methods
Georg Thimm and Emile Fiesler · 1995
Earlier work this paper cites.
Sparse Connection and Pruning in Large Dynamic Artificial Neural Networks
Nikko Ström · 1997
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O. Stanley and Risto Miikkulainen · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Multiplicative lstm for sequence modelling
Ben Krause, Liang Lu, Iain Murray, and Steve Renals · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
A survey on internet of things: Architecture, enabling technologies, security and privacy, and applications
J. Lin, W. Yu, N. Zhang, X. Yang, H. Zhang, and W. Zhao · 2017
Cited alongside, same era.
Variational Dropout Sparsifies Deep Neural Networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry P. Vetrov · 2017
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Sharan Narang, Greg Diamos, Shubho Sengupta, and Erich Elsen · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert A. Legenstein · 2018
Cited alongside, same era.
Learning sparse neural networks through
Diederik P. Kingma Christos Louizos, Max Welling · 2018
Estimation of energy consumption in machine learning
Eva García-Martín, Crefeda Faviola Rodrigues, Graham Riley, and Håkan Grahn · 2019
Later among the works it cites.
Streaming end-to-end speech recognition for mobile devices
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K. C. Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein · 2019
Later among the works it cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Cited alongside, same era.
Compression of end-to-end models
Ruoming Pang, Tara Sainath, Rohit Prabhavalkar, Suyog Gupta, Yonghui Wu, Shuyuan Zhang, and Chung-Cheng Chiu · 2018
Cited alongside, same era.
Data security and privacy-preserving in edge computing paradigm: Survey and open issues
J. Zhang, B. Chen, Y. Zhao, X. Cheng, and F. Hu · 2018
Cited alongside, same era.
To Prune, or Not to Prune: Exploring the Efficacy of Pruning for Model Compression
Michael Zhu and Suyog Gupta · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Augmenting self-attention with persistent memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Discovering neural wirings
Mitchell Wortsman, Ali Farhadi, and Mohammad Rastegari · 2019
Later among the works it cites.
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski · 2019
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity, 2020
Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi · 2020
Later among the works it cites.
Nvidia a100 tensor core gpu architecture, 2020
NVIDIA · 2020
Later among the works it cites.