Fetching the paper…
Reading the bibliography…
In Theory IIb we characterize with a mix of theory and experiments the optimization of deep convolutional networks by Stochastic Gradient Descent.
B. Gidas, “Blobal optimization via the Langevin equation,”
1985
Earlier work this paper cites.
S. Gelfand and S. Mitter, “Recursive stochastic algorithms for global optimization in
1991
Earlier work this paper cites.
Athena Scientific, Belmont, MA, 1996
D. P. Bertsekas and J. N. Tsitsiklis, · 1996
Earlier work this paper cites.
D. Bertsekas and J. Tsitsiklis, “Gradient Convergence in Gradient methods with Errors,”
2000
Earlier work this paper cites.
revised, oct 2012
L. Bottou, “Online algorithms and stochastic approximations,” in · 2012
Cited alongside, same era.
T. Poggio, H. Mhaskar, L. Rosasco, B. Miranda, and Q. Liao, “Why and when can deep - but not shallow - networks avoid the curse of dimensionality: a review,” tech. rep., MIT Center for Brains, Minds and Machines, 2016
2016
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima,” · 2016
Cited alongside, same era.
T. Poggio and Q. Liao, “Theory ii: Landscape of the empirical risk in deep learning,”
2017
Cited alongside, same era.
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio, “Sharp minima can generalize for deep nets,”
2017
Later among the works it cites.
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N.Golowich, and T. Poggio, “Musings on deep learning: Optimization properties of sgd,”
2017
Later among the works it cites.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…