Fetching the paper…
Reading the bibliography…
The deployment constraints in practical applications necessitate the pruning of large-scale deep learning models, i.e., promoting their weight sparsity.
“Information geonetry and alternating minimization procedures,”
Imre Csiszár, · 1984
Earlier work this paper cites.
“Optimal brain damage,”
Yann LeCun, John Denker, and Sara Solla, · 1989
Earlier work this paper cites.
“Pruning versus clipping in neural networks,”
Steven A Janowsky, · 1989
Earlier work this paper cites.
“Skeletonization: A technique for trimming the fat from a network via relevance assessment,”
Michael C Mozer and Paul Smolensky, · 1989
Earlier work this paper cites.
“Optimal brain damage,”
Yann LeCun, John S Denker, and Sara A Solla, · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork, · 1993
Earlier work this paper cites.
“A penalty function approach for solving bi-level linear programs,”
Douglas J White and G Anandalingam, · 1993
Earlier work this paper cites.
“Descent approaches for quadratic bilevel programming,”
Luis Vicente, Gilles Savard, and Joaquim Júdice, · 1994
Earlier work this paper cites.
“On bilevel programming, part i: general nonlinear cases,”
James E Falk and Jiming Liu, · 1995
Earlier work this paper cites.
“Learning multiple layers of features from tiny images,”
A. Krizhevsky and G. Hinton, · 2009
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
“Optimization with sparsity-inducing penalties,”
Francis Bach, Rodolphe Jenatton, Julien Mairal, Guillaume Obozinski, et al., · 2012
Earlier work this paper cites.
“Predicting parameters in deep learning,”
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando De Freitas, · 2013
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally, · 2015
Earlier work this paper cites.
“Learning both weights and connections for efficient neural network,”
Song Han, Jeff Pool, John Tran, and William Dally, · 2015
Earlier work this paper cites.
Uri Shaham, Yutaro Yamada, and Sahand Negahban, · 2015
Earlier work this paper cites.
“Tiny imagenet visual recognition challenge,”
Ya Le and Xuan Yang, · 2015
Earlier work this paper cites.
“Pruning convolutional neural networks for resource efficient inference,”
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz, · 2016
Earlier work this paper cites.
“Less is more: Towards compact cnns,”
Hao Zhou, Jose M Alvarez, and Fatih Porikli, · 2016
Earlier work this paper cites.
“Pruning filters for efficient convnets,”
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf, · 2016
Earlier work this paper cites.
Stephen Gould, Basura Fernando, Anoop Cherian, Peter Anderson, Rodrigo Santa Cruz, and Edison Guo, · 2016
Earlier work this paper cites.
“Learning structured sparsity in deep neural networks,”
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Exploring the granularity of sparsity in convolutional neural networks,”
Huizi Mao, Song Han, Jeff Pool, Wenshuo Li, Xingyu Liu, Yu Wang, and William J Dally, · 2017
Earlier work this paper cites.
“To prune, or not to prune: exploring the efficacy of pruning for model compression,”
Michael Zhu and Suyog Gupta, · 2017
Earlier work this paper cites.
“Learning efficient convolutional networks through network slimming,”
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang, · 2017
Earlier work this paper cites.
“Channel pruning for accelerating very deep neural networks,”
Yihui He, Xiangyu Zhang, and Jian Sun, · 2017
Earlier work this paper cites.
“Learning sparse neural networks through l _ 0 l\_0 regularization,”
Christos Louizos, Max Welling, and Diederik P Kingma, · 2017
Earlier work this paper cites.
“A first order method for solving convex bilevel optimization problems,”
Shoham Sabach and Shimrit Shtern, · 2017
Earlier work this paper cites.
“Forward and reverse gradient-based hyperparameter optimization,”
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil, · 2017
Earlier work this paper cites.
“Model-agnostic meta-learning for fast adaptation of deep networks,”
Chelsea Finn, Pieter Abbeel, and Sergey Levine, · 2017
Earlier work this paper cites.
“Amc: Automl for model compression and acceleration on mobile devices,”
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han, · 2018
Earlier work this paper cites.
“Admm-nn: An algorithm-hardware co-design framework of dnns using alternating direction method of multipliers,” 2018
Ao Ren, Tianyun Zhang, Shaokai Ye, Jiayu Li, Wenyao Xu, Xuehai Qian, Xue Lin, and Yanzhi Wang, · 2018
Earlier work this paper cites.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks,”
Jonathan Frankle and Michael Carbin, · 2018
Earlier work this paper cites.
“Snip: Single-shot network pruning based on connection sensitivity,”
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr, · 2018
Cited alongside, same era.
“Approximation methods for bilevel programming,”
Saeed Ghadimi and Mengdi Wang, · 2018
Cited alongside, same era.
“Darts: Differentiable architecture search,”
Hanxiao Liu, Karen Simonyan, and Yiming Yang, · 2018
Cited alongside, same era.
“A systematic dnn weight pruning framework using alternating direction method of multipliers,”
Tianyun Zhang, Shaokai Ye, Kaiqi Zhang, Jian Tang, Wujie Wen, Makan Fardad, and Yanzhi Wang, · 2018
Cited alongside, same era.
“Drawing early-bird tickets: Towards more efficient training of deep networks,”
“An image enhancing pattern-based sparsity for real-time inference on mobile devices,”
Xiaolong Ma, Wei Niu, Tianyun Zhang, Sijia Liu, Sheng Lin, Hongjia Li, Wujie Wen, Xiang Chen, Jian Tang, Kaisheng Ma, et al., · 2020
Later among the works it cites.
“Single shot structured pruning before training,”
Joost van Amersfoort, Milad Alizadeh, Sebastian Farquhar, Nicholas Lane, and Yarin Gal, · 2020
Later among the works it cites.
“A generic first-order algorithmic framework for bi-level programming beyond lower-level singleton,”
Risheng Liu, Pan Mu, Xiaoming Yuan, Shangzhi Zeng, and Jin Zhang, · 2020
Later among the works it cites.
“Improved bilevel model: Fast and optimal algorithm with theoretical guarantee,”
Junyi Li, Bin Gu, and Heng Huang, · 2020
Later among the works it cites.
“Bilevel optimization: Nonasymptotic analysis and faster algorithms,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haoran You, Chaojian Li, Pengfei Xu, Yonggan Fu, Yue Wang, Xiaohan Chen, Richard G Baraniuk, Zhangyang Wang, and Yingyan Lin, · 2019
Cited alongside, same era.
“An ADMM based framework for automl pipeline configuration,” 2019
Sijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf, Gregory Bramble, Horst Samulowitz, Dakuo Wang, Andrew Conn, and Alexander Gray, · 2019
Cited alongside, same era.
“Importance estimation for neural network pruning,”
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz, · 2019
Cited alongside, same era.
“The state of sparsity in deep neural networks,”
Trevor Gale, Erich Elsen, and Sara Hooker, · 2019
Cited alongside, same era.
“Truncated back-propagation for bilevel optimization,”
Amirreza Shaban, Ching-An Cheng, Nathan Hatch, and Byron Boots, · 2019
Cited alongside, same era.
“Meta-learning with implicit gradients,”
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine, · 2019
Cited alongside, same era.
“Research on distributed renewable energy transaction decision-making based on multi-agent bilevel cooperative reinforcement learning,”
Zhangyu Chen, Dong Liu, Xiaofei Wu, and Xiaochun Xu, · 2019
Cited alongside, same era.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
Kaiyi Ji, Junjie Yang, and Yingbin Liang, · 2020
Later among the works it cites.
Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang, · 2020
Later among the works it cites.
“On the iteration complexity of hypergradient computation,”
Riccardo Grazzi, Luca Franceschi, Massimiliano Pontil, and Saverio Salzo, · 2020
Later among the works it cites.
“Metapoison: Practical general-purpose clean-label data poisoning,”
W Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor, and Tom Goldstein, · 2020
Later among the works it cites.
“Dsa: More efficient budgeted pruning via differentiable sparsity allocation,”
Xuefei Ning, Tianchen Zhao, Wenshuo Li, Peng Lei, Yu Wang, and Huazhong Yang, · 2020
Later among the works it cites.
“A proximal iteratively reweighted approach for efficient network sparsification,”
Hao Wang, Xiangyu Yang, Yuanming Shi, and Jun Lin, · 2020
Later among the works it cites.
“On the decision boundaries of neural networks: A tropical geometry perspective,”
Motasem Alfarra, Adel Bibi, Hasan Hammoud, Mohamed Gaafar, and Bernard Ghanem, · 2020
Later among the works it cites.
“Deep learning: a statistical viewpoint,”
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin, · 2021
Later among the works it cites.
“A winning hand: Compressing deep networks can improve out-of-distribution robustness,”
James Diffenderfer, Brian Bartoldson, Shreya Chaganti, Jize Zhang, and Bhavya Kailkhura, · 2021
Later among the works it cites.
“The lottery tickets hypothesis for supervised and self-supervised pre-training in computer vision models,”
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, and Zhangyang Wang, · 2021
Later among the works it cites.
“Sanity checks for lottery tickets: Does your winning ticket really win the jackpot?,”
Xiaolong Ma, Geng Yuan, Xuan Shen, Tianlong Chen, Xuxi Chen, Xiaohan Chen, Ning Liu, Minghai Qin, Sijia Liu, Zhangyang Wang, et al., · 2021
Later among the works it cites.
“Efficient lottery ticket finding: Less data is more,”
Zhenyu Zhang, Xuxi Chen, Tianlong Chen, and Zhangyang Wang, · 2021
Later among the works it cites.
“Ac/dc: Alternating compressed/decompressed training of deep neural networks,”
Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh, · 2021
Later among the works it cites.
“Gdp: Stabilized neural network pruning via gates with differentiable polarization,”
Yi Guo, Huan Yuan, Jianchao Tan, Zhangyang Wang, Sen Yang, and Ji Liu, · 2021
Later among the works it cites.
“Differentiable network pruning for microcontrollers,”
Edgar Liberis and Nicholas D Lane, · 2021
Later among the works it cites.
“Rethinking bi-level optimization in neural architecture search: A gibbs sampling perspective,”
Chao Xue, Xiaoxing Wang, Junchi Yan, Yonggang Hu, Xiaokang Yang, and Kewei Sun, · 2021
Later among the works it cites.
“Effective sparsification of neural networks with global sparsity constraint,”
Xiao Zhou, Weizhong Zhang, Hang Xu, and Tong Zhang, · 2021
Later among the works it cites.
“Efficient lottery ticket finding: Less data is more,”
Zhenyu Zhang, Xuxi Chen, Tianlong Chen, and Zhangyang Wang, · 2021
Later among the works it cites.
“{GAN}s can play lottery tickets too,”
Xuxi Chen, Zhenyu Zhang, Yongduo Sui, and Tianlong Chen, · 2021
Later among the works it cites.
“Good students play big lottery better,”
Haoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You, Xiaohui Xie, and Zhangyang Wang, · 2021
Later among the works it cites.
“Playing lottery tickets with vision and language,”
Zhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen, Yu Cheng, Shuohang Wang, and Jingjing Liu, · 2021
Later among the works it cites.
“A unified lottery ticket hypothesis for graph neural networks,”
Tianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang, and Zhangyang Wang, · 2021
Later among the works it cites.
“Winning lottery tickets in deep generative models,” 2021
Neha Mukund Kalibhat, Yogesh Balaji, and Soheil Feizi, · 2021
Later among the works it cites.
“Ultra-data-efficient gan training: Drawing a lottery ticket first, then training it toughly,”
Tianlong Chen, Yu Cheng, Zhe Gan, Jingjing Liu, and Zhangyang Wang, · 2021
Later among the works it cites.
“Revisiting and advancing fast adversarial training through the lens of bi-level optimization,”
Yihua Zhang, Guanhuan Zhang, Prashant Khanduri, Mingyi Hong, Shiyu Chang, and Sijia Liu, · 2021
Later among the works it cites.
“Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,”
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste, · 2021
Later among the works it cites.
“Is the meta-learning idea able to improve the generalization of deep neural networks on the standard supervised learning?,”
Xiang Deng and Zhongfei Mark Zhang, · 2021
Later among the works it cites.
“Coarsening the granularity: Towards structurally sparse lottery tickets,”
Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang, and Zhangyang Wang, · 2022
Closest in time.
“Prospect pruning: Finding trainable weights at initialization using meta-gradients,”
Milad Alizadeh, Shyam A. Tailor, Luisa M Zintgraf, Joost van Amersfoort, Sebastian Farquhar, Nicholas Donald Lane, and Yarin Gal, · 2022
Closest in time.
“Paca: A pattern pruning algorithm and channel-fused high pe utilization accelerator for cnns,”
Jingyu Wang, Songming Yu, Zhuqing Yuan, Jinshan Yue, Zhe Yuan, Ruoyang Liu, Yanzhi Wang, Huazhong Yang, Xueqing Li, and Yongpan Liu, · 2022
Closest in time.
“Gradient-based bi-level optimization for deep learning: A survey,”
Can Chen, Xi Chen, Chen Ma, Zixuan Liu, and Xue Liu, · 2022
Closest in time.