Fetching the paper…
Reading the bibliography…
Deep neural networks often suffer from poor generalization due to complex and non-convex loss landscapes.
R. A. Fisher, “On the mathematical foundations of theoretical statistics,”
1922
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in
1937
Earlier work this paper cites.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,”
1989
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Simplifying neural nets by discovering flat minima,”
1994
Earlier work this paper cites.
L. K. Hansen, J. Larsen, and T. Fog, “Early stop criterion from the bootstrap ensemble,” in
1997
Earlier work this paper cites.
D. A. McAllester, “Pac-bayesian model averaging,” in
1999
Earlier work this paper cites.
B. Laurent and P. Massart, “Adaptive estimation of a quadratic functional by model selection,”
2000
Earlier work this paper cites.
K. Chellapilla, S. Puri, and P. Simard, “High performance convolutional neural networks for document processing,” in
2006
Earlier work this paper cites.
A. Krizhevsky
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
L. Bottou, “Large-scale machine learning with stochastic gradient descent,” in
2010
Earlier work this paper cites.
2012
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,”
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Martens and R. Grosse, “Optimizing neural networks with kronecker-factored approximate curvature,” in
2015
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,”
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Earlier work this paper cites.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” 2016
2016
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,”
2016
Earlier work this paper cites.
R. Grosse and J. Martens, “A kronecker-factored approximate fisher matrix for convolution layers,” in
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”
2017
Earlier work this paper cites.
T. DeVries and G. W. Taylor, “Improved regularization of convolutional neural networks with cutout,”
2017
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,”
2017
Cited alongside, same era.
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro, “Exploring generalization in deep learning,”
2017
Cited alongside, same era.
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio, “Sharp minima can generalize for deep nets,” in
2017
Cited alongside, same era.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska
2017
Cited alongside, same era.
U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen, “Rigging the lottery: Making all tickets winners,” in
2020
Later among the works it cites.
S. Jayakumar, R. Pascanu, J. Rae, S. Osindero, and E. Elsen, “Top-kast: Top-k always sparse training,”
2020
Later among the works it cites.
S. P. Singh and D. Alistarh, “Woodfisher: Efficient second-order approximation for neural network compression,”
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Q. Huang, K. Zhou, S. You, and U. Neumann, “Learning to prune filters in convolutional neural networks,” in
2018
Cited alongside, same era.
J. Kwon, J. Kim, H. Park, and I. K. Choi, “Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y.-L. Sung, V. Nair, and C. A. Raffel, “Training neural networks with fixed sparse masks,”
2021
Later among the works it cites.
S. Liu, T. Chen, X. Chen, Z. Atashgahi, L. Yin, H. Kou, L. Shen, M. Pechenizkiy, Z. Wang, and D. C. Mocanu, “Sparse training via boosting pruning plasticity with neuroregeneration,”
2021
Later among the works it cites.
S. Liu, T. Chen, X. Chen, L. Shen, D. C. Mocanu, Z. Wang, and M. Pechenizkiy, “The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Hubara, B. Chmiel, M. Island, R. Banner, J. Naor, and D. Soudry, “Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks,”
2021
Later among the works it cites.
J. Pool and C. Yu, “Channel permutations for n: m sparsity,”
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Andriushchenko and N. Flammarion, “Understanding sharpness-aware minimization,” 2021
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Chen, L. Shen, F. Zou, and W. Liu, “Towards practical adam: Non-convexity, convergence theory, and mini-batch acceleration,”
2022
Later among the works it cites.