Fetching the paper…
Reading the bibliography…
Modern neural network performance typically improves as model size increases.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Earlier work this paper cites.
A downsampled variant of imagenet as an alternative to the cifar datasets
Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter · 2017
Earlier work this paper cites.
Condensenet: An efficient densenet using learned group convolutions
Gao Huang, Shichen Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Modern neural networks generalize on small data sets
Matthew Olson, Abraham J. Wyner, and Richard Berk · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S. Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Later among the works it cites.
On the convex behavior of deep neural networks in relation to the layers’ width
Etai Littwin and L. Wolf · 2019
Later among the works it cites.
Fully learnable group convolution for acceleration of deep neural networks
Xijun Wang, Meina Kan, Shiguang Shan, and Xilin Chen · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
Efficient structured pruning and architecture searching for group convolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Ruizhe Zhao and Wayne Luk · 2019
Later among the works it cites.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2020
Closest in time.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2020
Closest in time.
Etai Littwin and L. Wolf · 2020
Closest in time.