Fetching the paper…
Reading the bibliography…
Overparameterized neural networks generalize well but are expensive to train.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
An algorithm for the machine calculation of complex fourier series
James W Cooley and John W Tukey · 1965
Earlier work this paper cites.
The diagonal decomposition technique applied to the dynamic programming solution of elliptic partial differential equations
DC Collins and ES Angel · 1971
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
Random butterfly transformations with applications in computational linear algebra
D Stott Parker · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Sparse coding in the primate cortex
Peter Foldiak · 2003
Earlier work this paper cites.
The penn treebank: an overview
Ann Taylor, Mitchell Marcus, and Beatrice Santorini · 2003
Earlier work this paper cites.
Stable signal recovery from incomplete and inaccurate measurements
Emmanuel J Candes, Justin K Romberg, and Terence Tao · 2006
Earlier work this paper cites.
Perturbed identity matrices have high rank: Proof and applications
Noga Alon · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright · 2011
Earlier work this paper cites.
Clustering partially observed graphs via convex optimization
Ali Jalali, Yudong Chen, Sujay Sanghavi, and Huan Xu · 2011
Earlier work this paper cites.
High dimensional low rank and sparse covariance matrix estimation via convex minimization
Xi Luo · 2011
Earlier work this paper cites.
CUDA Programming: A Developer’s Guide to Parallel Computing with GPUs
Shane Cook · 2012
Earlier work this paper cites.
Fast approximation of rotations and Hessians matrices
Michael Mathieu and Yann LeCun · 2014
Earlier work this paper cites.
Variable selection is hard
Dean Foster, Howard Karloff, and Justin Thaler · 2015
Earlier work this paper cites.
Butterfly factorization
Yingzhou Li, Haizhao Yang, Eileen R. Martin, Kenneth L. Ho, and Lexing Ying · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Jialin Pan · 2017
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Scott Gray, Alec Radford, and Diederik P Kingma · 2017
Earlier work this paper cites.
Tunable efficient unitary neural networks (EUNN) and their application to RNNs
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljacić · 2017
Earlier work this paper cites.
Runtime neural pruning
Ji Lin, Yongming Rao, Jiwen Lu, and Jie Zhou · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
A two-pronged progress in structured dense matrix vector multiplication
Christopher De Sa, Albert Gu, Rohan Puttagunta, Christopher Ré, and Atri Rudra · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Later among the works it cites.
Resprop: Reuse sparsified backpropagation
Negar Goli and Tor M. Aamodt · 2020
Later among the works it cites.
Sara Hooker · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Finding trainable sparse networks through neural tangent transfer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Quadrature-based features for kernel approximation
Marina Munkhoeva, Yermek Kapushev, Evgeny Burnaev, and Ivan Oseledets · 2018
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
Beidi Chen, Tharun Medini, James Farwell, Sameh Gobriel, Charlie Tai, and Anshumali Shrivastava · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Unifying orthogonal Monte Carlo methods
Krzysztof Choromanski, Mark Rowland, Wenyu Chen, and Adrian Weller · 2019
Cited alongside, same era.
Learning fast algorithms for linear transforms using butterfly factorizations
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 2019
Cited alongside, same era.
Tianlin Liu and Friedemann Zenke · 2020
Later among the works it cites.
Proving the lottery ticket hypothesis: Pruning is all you need
Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir · 2020
Later among the works it cites.
Logarithmic pruning is all you need
Laurent Orseau, Marcus Hutter, and Omar Rivasplata · 2020
Later among the works it cites.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Optimal lottery tickets via subsetsum: Logarithmic over-parameterization is sufficient
Ankit Pensia, Shashank Rajput, Alliot Nagle, Harit Vishwakarma, and Dimitris Papailiopoulos · 2020
Later among the works it cites.
Sparse weight activation training
Md Aamir Raihan and Tor M Aamodt · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M Rush · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel LK Yamins, and Surya Ganguli · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Later among the works it cites.
Butterfly transform: An efficient fft based neural architecture design
Keivan Alizadeh Vahid, Anish Prabhu, Ali Farhadi, and Mohammad Rastegari · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Later among the works it cites.
Lite transformer with long-short range attention
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Closest in time.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Closest in time.
Scatterbrain: Unifying sparse and low-rank attention
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher Ré · 2021
Closest in time.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Ž \́mathbf{i} · 2021
Closest in time.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, et al · 2021
Closest in time.
Nystromformer: A Nystrom-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan · 2021
Closest in time.
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Closest in time.
Cpm: A large-scale generative chinese pre-trained language model
Zhengyan Zhang, Xu Han, Hao Zhou, Pei Ke, Yuxian Gu, Deming Ye, Yujia Qin, Yusheng Su, Haozhe Ji, Jian Guan, et al · 2021
Closest in time.