Fetching the paper…
Reading the bibliography…
Neural network sparsification is a promising avenue to save computational time and memory costs, especially in an age where many successful AI models are becoming too large to na\"ively deploy on consumer hardware.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
A Practical Bayesian Framework for Backpropagation Networks
David J. C. MacKay · 1992
Earlier work this paper cites.
Bayesian learning via stochastic dynamics
Radford Neal · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Bayesian nonlinear modeling for the prediction competition
David JC MacKay et al · 1994
Earlier work this paper cites.
Probable networks and plausible predictions - a review of practical Bayesian methods for supervised neural networks
David John Cameron MacKay · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Occam's Razor
Carl Rasmussen and Zoubin Ghahramani · 2000
Earlier work this paper cites.
Sparse Bayesian learning and the relevance vector machine
Michael E Tipping · 2001
Earlier work this paper cites.
Information Theory, Inference & Learning Algorithms
David J. C. MacKay · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Sparse Bayesian learning for basis selection
David P Wipf and Bhaskar D Rao · 2004
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge, 2015
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Practical Gauss-Newton Optimisation for Deep Learning, 2017
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Earlier work this paper cites.
Net-trim: Convex pruning of deep neural networks with performance guarantee
Alireza Aghasi, Afshin Abdi, Nam Nguyen, and Justin Romberg · 2017
Cited alongside, same era.
Group sparse regularization for deep neural networks
Simone Scardapane, Danilo Comminiello, Amir Hussain, and Aurelio Uncini · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Wide Residual Networks, 2017
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A Scalable Laplace Approximation for Neural Networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Pruning and quantization for deep neural network acceleration: A survey
Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
A Gradient Flow Framework For Analyzing Network Pruning, 2021
Ekdeep Singh Lubana and Robert P. Dick · 2021
Later among the works it cites.
Efficient Neural Network Training via Forward and Backward Propagation Sparsification, 2021
Xiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao, and Tong Zhang · 2021
Later among the works it cites.
Only Train Once: A One-Shot Neural Network Training And Pruning Framework
Tianyi Chen, Bo Ji, DING Tianyu, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Sheng Yi, and Xiao Tu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Morphnet: Fast & simple resource-constrained structure learning of deep networks
Ariel Gordon, Elad Eban, Ofir Nachum, Bo Chen, Hao Wu, Tien-Ju Yang, and Edward Choi · 2018
Cited alongside, same era.
Fast approximate natural gradient descent in a kronecker factored eigenbasis
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, and Pascal Vincent · 2018
Cited alongside, same era.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip H. S. Torr · 2018
Cited alongside, same era.
Deep Rewiring: Training very sparse deep networks, 2018
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2018
Cited alongside, same era.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Fast Sparse ConvNets, 2019
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan · 2019
Cited alongside, same era.
MLP-Mixer: An all-MLP Architecture for Vision, 2021
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Robust Speech Recognition via Large-Scale Weak Supervision, 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
A review on TinyML: State-of-the-art and Prospects
Partha Pratim Ray · 2022
Later among the works it cites.
Challenges in Deploying Machine Learning: A Survey of Case Studies
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D. Lawrence · 2022
Later among the works it cites.
Adapting the Linearised Laplace Model Evidence for Modern Deep Learning
Javier Antorán, David Janz, James U Allingham, Erik Daxberger, Riccardo Rb Barbano, Eric Nalisnick, and José Miguel Hernández-Lobato · 2022
Later among the works it cites.
Winning the Lottery Ahead of Time: Efficient Early Network Pruning, 2022
John Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler, Morgane Ayle, and Stephan Günnemann · 2022
Later among the works it cites.
Invariance learning in deep neural networks with differentiable Laplace approximations
Alexander Immer, Tycho van der Ouderaa, Gunnar Rätsch, Vincent Fortuin, and Mark van der Wilk · 2022
Later among the works it cites.
Better uncertainty calibration via proper scores for classification and beyond
Sebastian Gruber and Florian Buettner · 2022
Later among the works it cites.
A Simple and Effective Pruning approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter · 2023
Later among the works it cites.
Depgraph: Towards any structural pruning
Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang · 2023
Later among the works it cites.
Stochastic marginal likelihood gradients using neural tangent kernels
Alexander Immer, Tycho FA Van Der Ouderaa, Mark Van Der Wilk, Gunnar Ratsch, and Bernhard Schölkopf · 2023
Later among the works it cites.
Online Laplace model selection revisited
Jihao Andreas Lin, Javier Antorán, and José Miguel Hernández-Lobato · 2023
Later among the works it cites.
Otov2: Automatic, generic, user-friendly
Tianyi Chen, Luming Liang, DING Tianyu, Zhihui Zhu, and Ilya Zharkov · 2023
Later among the works it cites.
ASDL: A Unified Interface for Gradient Preconditioning in PyTorch, 2023
Kazuki Osawa, Satoki Ishikawa, Rio Yokota, Shigang Li, and Torsten Hoefler · 2023
Later among the works it cites.
Probabilistic pathway-based multimodal factor analysis
Alexander Immer, Stefan G Stark, Francis Jacob, Ximena Bonilla, Tinu Thomas, Andre Kahles, Sandra Goetze, Emanuela S Milani, and Bernd Wollscheid · 2024
Closest in time.
Improving Neural Additive Models with Bayesian Principles
Kouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Ratsch, and Vincent Fortuin · 2024
Closest in time.