Fetching the paper…
Reading the bibliography…
Dynamic Sparse Training (DST) methods achieve state-of-the-art results in sparse neural network training, matching the generalization of dense models while enabling sparse training and inference.
Rigging the Lottery: Making All Tickets Winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 1911
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Torchvision: Pytorch’s computer vision library
Maintainers and Contributors · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
pytorch-cifar, 2017
Kuang Liu · 2017
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Random Erasing Data Augmentation, November 2017
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Earlier work this paper cites.
Deep Rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2018
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour, April 2018
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
AutoAugment: Learning Augmentation Policies from Data, April 2019
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2019
Earlier work this paper cites.
Sparse Networks from Scratch: Faster Training without Losing Performance
Tim Dettmers and Luke Zettlemoyer · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, May 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Network pruning via transformable architecture search
Xuanyi Dong and Yi Yang · 2019
Earlier work this paper cites.
Filter pruning via geometric median for deep convolutional neural networks acceleration
Yang He, Ping Liu, Ziwei Wang, Zhilan Hu, and Yi Yang · 2019
Cited alongside, same era.
Searching for mobilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam · 2019
Cited alongside, same era.
Accelerating deep learning inference via freezing
Adarsh Kumar, Arjun Balasubramanian, Shivaram Venkataraman, and Aditya Akella · 2019
Cited alongside, same era.
Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing
En Li, Liekang Zeng, Zhi Zhou, and Xu Chen · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang · 2019
Cited alongside, same era.
Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
How well do sparse imagenet models transfer?
Eugenia Iofinova, Alexandra Peste, Mark Kurtz, and Dan Alistarh · 2021
Later among the works it cites.
Sparse is enough in scaling transformers
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Later among the works it cites.
The future is log-gaussian: Resnets and their infinite-depth-and-width limit at initialization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhonghui You, Kun Yan, Jinmian Ye, Meng Ma, and Ping Wang · 2019
Cited alongside, same era.
Autoslim: Towards one-shot architecture search for channel numbers
Jiahui Yu and Thomas Huang · 2019
Cited alongside, same era.
CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features
Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Youngjoon Yoo, and Junsuk Choe · 2019
Cited alongside, same era.
Towards efficient model compression via learned global ranking
Ting-Wu Chin, Ruizhou Ding, Cha Zhang, and Diana Marculescu · 2020
Cited alongside, same era.
RandAugment: Practical Automated Data Augmentation with a Reduced Search Space
Ekin Dogus Cubuk, Barret Zoph, Jon Shlens, and Quoc Le · 2020
Cited alongside, same era.
Fast sparse convnets
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan · 2020
Cited alongside, same era.
Sparse gpu kernels for deep learning, 2020
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen · 2020
Cited alongside, same era.
Mufan Li, Mihai Nica, and Daniel M. Roy · 2021
Later among the works it cites.
Accelerating Sparse Deep Neural Networks, April 2021
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Later among the works it cites.
Deepsparse engine: Sparsity-aware deep learning inference runtime for CPUs, 2021
Neural Magic · 2021
Later among the works it cites.
Locally free weight sharing for network width search
Xiu Su, Shan You, Tao Huang, Fei Wang, Chen Qian, Changshui Zhang, and Chang Xu · 2021
Later among the works it cites.
Chip: Channel independence-based pruning for compact neural networks
Yang Sui, Miao Yin, Yi Xie, Huy Phan, Saman Aliari Zonouz, and Bo Yuan · 2021
Later among the works it cites.
MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
Geng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li, Zhenglun Kong, Ning Liu, Yifan Gong, Zheng Zhan, Chaoyang He, Qing Jin, Siyue Wang, Minghai Qin, Bin Ren, Yanzhi Wang, Sijia Liu, and Xue Lin · 2021
Later among the works it cites.
Learning n:m fine-grained structured sparse neural networks from scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li · 2021
Later among the works it cites.
Towards Structured Dynamic Sparse Pre-Training of BERT
Anastasia S. D. Dietrich, Frithjof Gressmann, Douglas Orr, Ivan Chelombiev, Daniel Justus, and Carlo Luschi · 2022
Later among the works it cites.
Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win - AAAI 2022 Poster, February 2022
Utku Evci, Yani Ioannou, cem Keskin, and Yann Dauphin · 2022
Later among the works it cites.
FSCNN: A Fast Sparse Convolution Neural Network Inference System, December 2022
Bo Ji and Tianyi Chen · 2022
Later among the works it cites.
Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse Training
Peng Jiang, Lihan Hu, and Shihui Song · 2022
Later among the works it cites.
Merlin HugeCTR: GPU-accelerated Recommender System Training and Inference
Zehuan Wang, Yingcan Wei, Minseok Lee, Matthias Langer, Fan Yu, Jie Liu, Shijie Liu, Daniel G. Abel, Xu Guo, Jianbing Dong, Ji Shi, and Kunlun Li · 2022
Later among the works it cites.
Get More at Once: Alternating Sparse Training with Gradient Correction
Li Yang, Jian Meng, Jae-sun Seo, and Deliang Fan · 2022
Later among the works it cites.
Erik Schultheis and Rohit Babbar · 2023
Closest in time.
Dynamic Sparsity Is Channel-Level Sparsity Learner
Lu Yin, Gen Li, Meng Fang, Li Shen, Tianjin Huang, Zhangyang Wang, Vlado Menkovski, Xiaolong Ma, Mykola Pechenizkiy, and Shiwei Liu · 2023
Closest in time.
mixup: Beyond Empirical Risk Minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2023
Closest in time.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H. Nguyen, Madeleine Gibescu, and Antonio Liotta · 2041
Closest in time.