Fetching the paper…
Reading the bibliography…
Despite impressive performance, deep neural networks require significant memory and computation costs, prohibiting their application in resource-constrained scenarios.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Variance reduction in stochastic gradient Langevin dynamics
Kumar Avinava Dubey, Sashank J Reddi, Sinead A Williamson, Barnabas Poczos, Alexander J Smola, and Eric P Xing · 2016
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alex Smola · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Earlier work this paper cites.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2018
Earlier work this paper cites.
Subsampled stochastic variance-reduced gradient Langevin dynamics
Difan Zou, Pan Xu, and Quanquan Gu · 2018
Earlier work this paper cites.
On the theory of variance reduction for stochastic gradient Monte Carlo
Niladri Chatterji, Nicolas Flammarion, Yian Ma, Peter Bartlett, and Michael Jordan · 2018
Earlier work this paper cites.
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal Mian · 2018
Earlier work this paper cites.
VR-SGD: A simple stochastic variance reduction method for machine learning
Fanhua Shang, Kaiwen Zhou, Hongying Liu, James Cheng, Ivor W Tsang, Lijun Zhang, Dacheng Tao, and Licheng Jiao · 2018
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Cited alongside, same era.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer · 2019
Cited alongside, same era.
Autoprune: Automatic network pruning by regularizing auxiliary parameters
Xia Xiao, Zigeng Wang, and Sanguthevar Rajasekaran · 2019
Cited alongside, same era.
Adversarial robustness vs. model compression, or both?
Shaokai Ye, Kaidi Xu, Sijia Liu, Hao Cheng, Jan-Henrik Lambrechts, Huan Zhang, Aojun Zhou, Kaisheng Ma, Yanzhi Wang, and Xue Lin · 2019
Cited alongside, same era.
A convergence analysis for a class of practical variance-reduction stochastic gradient MCMC
Changyou Chen, Wenlin Wang, Yizhe Zhang, Qinliang Su, and Lawrence Carin · 2019
Cited alongside, same era.
Resource-efficient deep neural networks for automotive radar interference mitigation
Johanna Rock, Wolfgang Roth, Mate Toth, Paul Meissner, and Franz Pernkopf · 2021
Later among the works it cites.
Optimal sensor channel selection for resource-efficient deep activity recognition
Clayton Frederick Souza Leite and Yu Xiao · 2021
Later among the works it cites.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
DNR: A tunable robust pruning framework through dynamic network rewiring of DNNs
Souvik Kundu, Mahdi Nazemi, Peter A Beerel, and Massoud Pedram · 2021
Later among the works it cites.
Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang · 2019
Cited alongside, same era.
Robust sparse regularization: Simultaneously optimizing neural network robustness and compactness
Adnan Siraj Rakin, Zhezhi He, Li Yang, Yanzhi Wang, Liqiang Wang, and Deliang Fan · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex SGD
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
Control variates for stochastic gradient MCMC
Jack Baker, Paul Fearnhead, Emily B Fox, and Christopher Nemeth · 2019
Cited alongside, same era.
Theoretically principled trade-off between robustness and accuracy
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Pixelated butterfly: Simple and efficient sparse training for neural network models
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant Jayakumar, Razvan Pascanu, Peter E Latham, and Yee Teh · 2021
Later among the works it cites.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
Hongwu Peng, Shaoyi Huang, Tong Geng, Ang Li, Weiwen Jiang, Hang Liu, Shusen Wang, and Caiwen Ding · 2021
Later among the works it cites.
Accelerating framework of transformer by hardware design and model compression co-optimization
Panjie Qi, Edwin Hsing-Mean Sha, Qingfeng Zhuge, Hongwu Peng, Shaoyi Huang, Zhenglun Kong, Yuhong Song, and Bingbing Li · 2021
Later among the works it cites.
[reproducibility report] rigging the lottery: Making all tickets winners
Varun Sundar and Rajat Vadiraj Dwaraknath · 2021
Later among the works it cites.
Implicit sparse regularization: The impact of depth and early stopping
Jiangyuan Li, Thanh Nguyen, Chinmay Hegde, and Ka Wai Wong · 2021
Later among the works it cites.
Comprehensive graph gradual pruning for sparse training in graph neural networks
Chuang Liu, Xueqi Ma, Yinbing Zhan, Liang Ding, Dapeng Tao, Bo Du, Wenbin Hu, and Danilo Mandic · 2022
Later among the works it cites.
The state of sparse training in deep reinforcement learning
Laura Graesser, Utku Evci, Erich Elsen, and Pablo Samuel Castro · 2022
Later among the works it cites.
Dynamic sparse training via balancing the exploration-exploitation trade-off
Shaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng, Yue Sun, Mimi Xie, and Caiwen Ding · 2022
Later among the works it cites.
Optimization with access to auxiliary information
El Mahdi Chayti and Sai Praneeth Karimireddy · 2022
Later among the works it cites.
Gradient flow in sparse neural networks and how lottery tickets win
Utku Evci, Yani Ioannou, Cem Keskin, and Yann Dauphin · 2022
Later among the works it cites.
Calibrating the rigged lottery: Making all tickets reliable
Bowen Lei, Ruqi Zhang, Dongkuan Xu, and Bani Mallick · 2023
Closest in time.
Neurogenesis dynamics-inspired spiking neural network training acceleration
Shaoyi Huang, Haowen Fang, Kaleel Mahmood, Bowen Lei, Nuo Xu, Bin Lei, Yue Sun, Dongkuan Xu, Wujie Wen, and Caiwen Ding · 2023
Closest in time.
Sparseprop: efficient sparse backpropagation for faster training of neural networks at the edge
Mahdi Nikdan, Tommaso Pegolotti, Eugenia Iofinova, Eldar Kurtic, and Dan Alistarh · 2023
Closest in time.
Efficient n: M sparse dnn training using algorithm, architecture, and dataflow co-design
Chao Fang, Wei Sun, Aojun Zhou, and Zhongfeng Wang · 2023
Closest in time.
Accelerating dataset distillation via model augmentation
Lei Zhang, Jie Zhang, Bowen Lei, Subhabrata Mukherjee, Xiang Pan, Bo Zhao, Caiwen Ding, Yao Li, and Dongkuan Xu · 2023
Closest in time.
Rethinking data distillation: Do not overlook calibration
Dongyao Zhu, Bowen Lei, Jie Zhang, Yanbo Fang, Yiqun Xie, Ruqi Zhang, and Dongkuan Xu · 2023
Closest in time.
Gradient matching for categorical data distillation in ctr prediction
Cheng Wang, Jiacheng Sun, Zhenhua Dong, Ruixuan Li, and Rui Zhang · 2023
Closest in time.
More is less: inducing sparsity via overparameterization
Hung-Hsu Chou, Johannes Maly, and Holger Rauhut · 2023
Closest in time.