Fetching the paper…
Reading the bibliography…
Recently, sparse training methods have started to be established as a de facto approach for training and inference efficiency in artificial neural networks.
Hesham Mostafa and Xin Wang · 1902
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On random graphs i
P Erdös and A Rényi · 1959
Earlier work this paper cites.
Optimal brain damage
Y. LeCun, J. Denker, and S. Solla · 1989
Earlier work this paper cites.
Using relevance to reduce network size automatically
Michael C. Mozer and Paul Smolensky · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Babak Hassibi, David G.Stork, and Gregory Wolff · 1993
Earlier work this paper cites.
6. the gene expression omnibus (geo): A gene expression and hybridization repository
Ron Edgar and Alex Lash · 2002
Earlier work this paper cites.
Design of experiments for the nips 2003 variable selection benchmark
I. Guyon · 2003
Earlier work this paper cites.
Result analysis of the nips 2003 feature selection challenge
Isabelle Guyon, Steve Gunn, Asa Ben-Hur, and Gideon Dror · 2003
Earlier work this paper cites.
The architecture of complex weighted networks
Alain Barrat, Marc Barthelemy, Romualdo Pastor-Satorras, and Alessandro Vespignani · 2004
Earlier work this paper cites.
Topological insights into sparse neural networks
S. Liu, Tim van der Lee, A. Yaman, Zahra Atashgahi, Davide Ferraro, Ghada Sokar, M. Pechenizkiy, and D. Mocanu · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Topology of a Neural Network , pp. 988–989
Risto Miikkulainen · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Graph Algorithms in the Language of Linear Algebra
Jeremy Kepner and John Gilbert · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, Christopher Ré, Stephen J. Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Predicting parameters in deep learning, 2013
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando de Freitas · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas · 2013
Earlier work this paper cites.
Network hubs in the human brain
Martijn P. van den Heuvel and Olaf Sporns · 2013
Earlier work this paper cites.
Revisiting asynchronous linear solvers: Provable convergence rate through randomization
H. Avron, Alex Druinsky, and A. Gupta · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks, 2015
Song Han, Jeff Pool, John Tran, and William J. Dally · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Numba: A llvm-based python jit compiler
Siu Kwan Lam, Antoine Pitrou, and Stanley Seibert · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Y. Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Y. Li, and J. Liu · 2015
Earlier work this paper cites.
Sparse convolutional neural networks
B. Liu, M. Wang, H. Foroosh, M. Tappen, and Marianna Pensky · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Suraj Srinivas and R. Venkatesh Babu · 2015
Earlier work this paper cites.
Network science
Albert-László Barabási and Márton Pósfai · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and W. Dally · 2016
Earlier work this paper cites.
Deep learning with s-shaped rectified linear activation units
X. Jin, Chunyan Xu, Jiashi Feng, Yunchao Wei, Junjun Xiong, and S. Yan · 2016
Cited alongside, same era.
Chapter 4 - deep learning and its parallelization
X. Li, G. Zhang, K. Li, and W. Zheng · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and C. Ré · 2016
Cited alongside, same era.
A topological insight into restricted boltzmann machines
Decebal Constantin Mocanu, Elena Mocanu, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2016
Cited alongside, same era.
Parallel sgd: When does averaging help?, 2016
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
How important is a neuron?, 2019
K. Dhamdhere, M. Sundararajan, and Qiqi Yan · 2019
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2019
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, A. Doucet, and J. Rousseau · 2019
Later among the works it cites.
Radix-net: Structured sparse matrices for deep neural networks
Jeremy Kepner and Ryan Robinett · 2019
Later among the works it cites.
Sparse deep neural network graph challenge
Jeremy Kepner, Simon Alford, Vijay N. Gadepally, Michael Jones, Lauren Milechin, Ryan A. Robinett, and Siddharth Samsi · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
An mpi-based python framework for distributed training with keras, 2017
Dustin Anderson, Jean-Roch Vlimant, and Maria Spiropulu · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2017
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
N. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, Raminder Bajwa, S. Bates, Suresh Bhatia, Nan Boden, Al Borchers, R. Boyle, P. Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, M. Daley, M. Dau, J. Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Ho, Doug Hogberg, John Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, Diemthu Le, C. Leary, Z. Liu, K. Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, K. Miller, R. Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, A. Phelps, J. Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, J. Souter, Dan A. Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, B. Tian, H. Toma, Erick Tuttle, V. Vasudevan, R. Walter, Walter Wang, E. Wilcox, and D. H. Yoon · 2017
Cited alongside, same era.
Enabling massive deep neural networks with the graphblas
J. Kepner, M. Kumar, J. Moreira, P. Pattnaik, M. Serrano, and H. Tufo · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H. McMahan, Eider Moore, D. Ramage, S. Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
Network computations in artificial intelligence
D.C. Mocanu · 2017
Cited alongside, same era.
Later among the works it cites.
Snip: Single-shot network pruning based on connection sensitivity, 2019
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr · 2019
Later among the works it cites.
Local sgd converges fast and communicates little, 2019
Sebastian U. Stich · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, S. Yang, and Shenghuo Zhu · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and J. Yosinski · 2019
Later among the works it cites.
Bigmpi4py: Python module for parallelization of big data objects
Alex M. Ascension and Marcos J. Araúzo-Bravo · 2020
Later among the works it cites.
Activation function impact on sparse neural networks
A. Dubowski · 2020
Later among the works it cites.
Gradient flow in sparse neural networks and how lottery tickets win
Utku Evci, Yani A. Ioannou, Cem Keskin, and Yann Dauphin · 2020
Later among the works it cites.
Sparse gpu kernels for deep learning
Trevor Gale, Matei A. Zaharia, Cliff Young, and Erich Elsen · 2020
Later among the works it cites.
Stochastic weight averaging in parallel: Large-batch training that generalizes well, 2020
Vipul Gupta, Santiago Akle Serrano, and Dennis DeCoste · 2020
Later among the works it cites.
Studying the effects of hashing of sparse deep neural networks on data and model parallelisms
M. Hasanzadeh-Mofrad, R. Melhem, M. Y. Ahmad, and Mohammad Hammoud · 2020
Later among the works it cites.
At-scale sparse deep neural network inference with efficient gpu implementation
M. Hidayetoğlu, C. Pearson, V. S. Mailthody, E. Ebrahimi, J. Xiong, R. Nagi, and W. m. Hwu · 2020
Later among the works it cites.
Top-kast: Top-k always sparse training
Siddhant Jayakumar, Razvan Pascanu, Jack Rae, Simon Osindero, and Erich Elsen · 2020
Later among the works it cites.
Accelerating sparsity in the nvidia ampere architecture, 2020
Jeff Pool · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity
Aditya Kusupati, V. Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham M. Kakade, and Ali Farhadi · 2020
Later among the works it cites.
A signal propagation perspective for pruning neural networks at initialization, 2020
Namhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, and Philip H. S. Torr · 2020
Later among the works it cites.
Don’t use large mini-batches, use local sgd, 2020
Tao Lin, S. Stich, and M. Jaggi · 2020
Later among the works it cites.
Combinatorial tiling for sparse neural networks
F. Pawłowski, R. Bisseling, B. Uçar, and A. Yzelman · 2020
Later among the works it cites.
Sparse weight activation training
Md Aamir Raihan and Tor Aamodt · 2020
Later among the works it cites.
Data parallel large sparse deep neural network on gpu
N. S. Sattar and Shaikh Anfuzzaman · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow, 2020
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
Asynchronous stochastic gradient descent with delay compensation, 2020
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2020
Later among the works it cites.
Keep the gradients flowing: Using gradient flow to study sparse network optimization
Kale ab Tessera, Sara Hooker, and Benjamin Rosman · 2021
Closest in time.
S. Liu, D. Mocanu, Yulong Pei, and Mykola Pechenizkiy · 2021
Closest in time.
Accelerating distributed inference of sparse deep neural networks via mitigating the straggler effect
Yousuf Ahmad Mohammad Hasanzadeh Mofrad, Rami Melhem and Mohammad Hammoud · 2021
Closest in time.
Growing efficient deep networks by structured continuous sparsification
Xin Yuan, Pedro H. P. Savarese, and Michael Maire · 2021
Closest in time.
Searching for exotic particles in high-energy physics with deep learning
P. Baldi, P. Sadowski, and D. Whiteson · 2041
Closest in time.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2041
Closest in time.