Fetching the paper…
Reading the bibliography…
Obtaining versions of deep neural networks that are both highly-accurate and highly-sparse is one of the main challenges in the area of model compression, and several high-performance pruning techniques have been investigated by the community.
Praktische verfahren der gleichungsauflösung
Richard von Mises and Hilda Pollaczek-Geiringer · 1929
Earlier work this paper cites.
Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories
Fei-Fei Li, R. Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush · 2005
Earlier work this paper cites.
The Caltech 256
Gregory Griffin, Alexander D. Holub, and Pietro Perona · 2006
Earlier work this paper cites.
A visual vocabulary for flower classification
Maria-Elena Nilsback and Andrew Zisserman · 2006
Earlier work this paper cites.
Iterative thresholding for sparse approximations
Thomas Blumensath and Mike E Davies · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Birdsnap: Large-scale fine-grained visual categorization of birds
Thomas Berg, Jiongxin Liu, Seung Woo Lee, Michelle L. Alexander, David W. Jacobs, and Peter N. Belhumeur · 2014
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J Dally · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Jialin Pan · 2017
Earlier work this paper cites.
MobileNets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Cited alongside, same era.
Efficient inference with TensorRT
Han Vanholder · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Keep the gradients flowing: Using gradient flow to study sparse network optimization, 2021
Kale ab Tessera, Sara Hooker, and Benjamin Rosman · 2021
Later among the works it cites.
M-FAC: Efficient matrix-free approximations of second-order information
Elias Frantar, Eldar Kurtic, and Dan Alistarh · 2021
Later among the works it cites.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
Sparse progressive distillation: Resolving overfitting under pretrain-and-finetune paradigm
Shaoyi Huang, Dongkuan Xu, Ian EH Yen, Yijue Wang, Sung-En Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Sanguthevar Rajasekaran, Hang Liu, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
pytorch-hessian-eigenthings: efficient pytorch hessian eigendecomposition, 2018
Noah Golmant, Zhewei Yao, Amir Gholami, Michael Mahoney, and Joseph Gonzalez · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Siddhant M Jayakumar, Razvan Pascanu, Jack W Rae, Simon Osindero, and Erich Elsen · 2021
Later among the works it cites.
Lost in Pruning: The Effects of Pruning Neural Networks beyond Test Accuracy
Lucas Liebenwein, Cenk Baykal, Brandon Carter, David Gifford, and Daniela Rus · 2021
Later among the works it cites.
Do we actually need dense over-parameterization? in-time over-parameterization in sparse training
Shiwei Liu, Lu Yin, Decebal Constantin Mocanu, and Mykola Pechenizkiy · 2021
Later among the works it cites.
AC/DC: Alternating compressed/decompressed training of deep neural networks
Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh · 2021
Later among the works it cites.
Winning the lottery with continuous sparsification
Pedro Savarese, Hugo Silva, and Michael Maire · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant Jayakumar, Razvan Pascanu, Peter Latham, and Yee Teh · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models
Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability, 2022
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2022
Later among the works it cites.
Gradmax: Growing neural networks using gradient information
Utku Evci, Max Vladymyrov, Thomas Unterthiner, Bart van Merrienboer, and Fabian Pedregosa · 2022
Later among the works it cites.
How well do sparse imagenet models transfer?
Eugenia Iofinova, Alexandra Peste, and Dan Alistarh · 2022
Later among the works it cites.
Training your sparse neural network better with any mask
Ajay Jaiswal, Haoyu Ma, Tianlong Chen, Ying Ding, and Zhangyang Wang · 2022
Later among the works it cites.
Gmp*: Well-tuned global magnitude pruning can outperform most bert-pruning methods
Eldar Kurtic and Dan Alistarh · 2022
Later among the works it cites.
The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models
Eldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar, Mark Kurtz, Benjamin Fineran, Michael Goin, and Dan Alistarh · 2022
Later among the works it cites.
Sparse training via boosting pruning plasticity with neuroregeneration
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi, Lu Yin, Huanyu Kou, Li Shen, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu · 2022
Later among the works it cites.
The DeepSparse Inference Engine
NeuralMagic · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance
Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2022
Later among the works it cites.
Bias in pruned vision models: In-depth analysis and countermeasures
Eugenia Iofinova, Alexandra Peste, and Dan Alistarh · 2023
Closest in time.
Step: Learning n:m structured sparsity masks from scratch with precondition, 2023
Yucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov, Christopher De Sa, and Amir Yazdanbakhsh · 2023
Closest in time.
Sparseprop: Efficient sparse propagation for faster training of neural networks
Mahdi Nikdan, Tomasso Pegolotti, Eugenia Iofinova, Eldar Kurtic, and Dan Alistarh · 2023
Closest in time.
Are straight-through gradients and soft-thresholding all you need for sparse training?
Antoine Vanderschueren and Christophe De Vleeschouwer · 2023
Closest in time.
Dynamic sparsity is channel-level sparsity learner
Lu Yin, Gen Li, Meng Fang, Li Shen, Tianjin Huang, Zhangyang Wang, Vlado Menkovski, Xiaolong Ma, Mykola Pechenizkiy, and Shiwei Liu · 2023
Closest in time.