Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is the workhorse algorithm of deep learning technology.
A. E. Hoerl and R. W. Kennard, Technometrics 12
1970
Earlier work this paper cites.
J. R. Fienup, Applied optics 21
1982
Earlier work this paper cites.
M. Mézard, G. Parisi, and M. A. Virasoro, Spin glass theory and beyond (World Scientific, Singapore, 1987)
1987
Earlier work this paper cites.
H. Sompolinsky, A. Crisanti, and H.-J. Sommers, Physical review letters 61
1988
Earlier work this paper cites.
V. Vapnik, in Advances in Neural Information Processing Systems , edited by J. Moody, S. Hanson, and R. P. Lippmann (Morgan-Kaufmann, 1992), vol. 4
1992
Earlier work this paper cites.
B. E. Boser, I. M. Guyon, and V. N. Vapnik, in Proceedings of the Fifth Annual Workshop on Computational Learning Theory (Association for Computing Machinery, New York, NY, USA, 1992), COLT ’92, p. 144–152, ISBN 089791497X, URL https://doi.org/10.1145/130385.130401
1992
Earlier work this paper cites.
L. F. Cugliandolo and J. Kurchan, Physical Review Letters 71
1993
Earlier work this paper cites.
A. Georges, G. Kotliar, W. Krauth, and M. J. Rozenberg, Reviews of Modern Physics 68
1996
Earlier work this paper cites.
S. Bös and M. Opper, in Advances in Neural Information Processing Systems (1997), pp. 141–147
1997
Earlier work this paper cites.
J. Kurchan, Jamming and Rheology: Constrained Dynamics on Microscopic and Macroscopic Scales 72
1997
Earlier work this paper cites.
L. Bottou (1999)
1999
Earlier work this paper cites.
Coolen, A. C. C., Saad, D., and Xiong, Yuan-Sheng, Europhys. Lett. 51
2000
Earlier work this paper cites.
E. Abbe and C. Sandon, Poly-time universality and limitations of deep learning (2020), eprint 2001.02992
2001
Earlier work this paper cites.
F. Mignacco, F. Krzakala, Y. M. Lu, and L. Zdeborová, arXiv preprint arXiv:2002.11544 (2020b)
2002
Earlier work this paper cites.
D. Loi, S. Mossa, and L. F. Cugliandolo, Physical Review E 77
2008
Earlier work this paper cites.
D. Saad, On-line learning in neural networks , vol. 17 (Cambridge University Press, 2009)
2009
Earlier work this paper cites.
H. Xu, C. Caramanis, and S. Mannor, Journal of machine learning research 10
2009
Earlier work this paper cites.
L. F. Cugliandolo, Journal of Physics A: Mathematical and Theoretical 44
2011
Earlier work this paper cites.
Y. Bengio, A. C. Courville, and P. Vincent, IEEE Transactions on Pattern Analysis and Machine Intelligence 35
2013
Cited alongside, same era.
A. M. Saxe, J. L. McClelland, and S. Ganguli, arXiv preprint arXiv:1312.6120 (2013)
2013
Cited alongside, same era.
M. C. Marchetti, J.-F. Joanny, S. Ramaswamy, T. B. Liverpool, J. Prost, M. Rao, and R. A. Simha, Reviews of Modern Physics 85
2013
Cited alongside, same era.
L. Berthier and J. Kurchan, Nature Physics 9
2013
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, nature 521
2015
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, ArXiv abs/1611.03530
B. Wu, W. Chen, Y. Fan, Y. Zhang, J. Hou, J. Liu, and T. Zhang, IEEE Access 7
2019
Later among the works it cites.
W. Hu, C. J. Li, L. Li, and J.-G. Liu, Annals of Mathematical Sciences and Applications 4
2019
Later among the works it cites.
U. Simsekli, L. Sagun, and M. Gurbuzbalaban, in International Conference on Machine Learning (PMLR, 2019), pp. 5827–5837
2019
Later among the works it cites.
S. Goldt, M. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, in Advances in Neural Information Processing Systems (2019), pp. 6979–6989
2019
Later among the works it cites.
S. Franz, S. Hwang, and P. Urbani, Physical review letters 123
2019
Later among the works it cites.
L. Zdeborová, Nature Physics 16
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, On large-batch training for deep learning: Generalization gap and sharp minima (2017), eprint 1609.04836
2017
Cited alongside, same era.
Q. Li, C. Tai, and E. Weinan, in International Conference on Machine Learning (PMLR, 2017), pp. 2101–2110
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Franz, G. Parisi, M. Sevelev, P. Urbani, and F. Zamponi, SciPost Physics 2
2017
Cited alongside, same era.
S. Yaida, arXiv preprint arXiv:1810.00004 (2018)
2018
Cited alongside, same era.
G. M. Rotskoff and E. Vanden-Eijnden, arXiv preprint arXiv:1805.00915 (2018)
2018
Cited alongside, same era.
Later among the works it cites.
J. Z. HaoChen, C. Wei, J. D. Lee, and T. Ma, arXiv preprint arXiv:2006.08680 (2020)
2020
Later among the works it cites.
X. Cheng, D. Yin, P. Bartlett, and M. Jordan, in International Conference on Machine Learning (PMLR, 2020), pp. 1810–1819
2020
Later among the works it cites.
F. Mignacco, F. Krzakala, P. Urbani, and L. Zdeborová, NIPS 2020 (2020a)
2020
Later among the works it cites.
G. Parisi, P. Urbani, and F. Zamponi, Theory of Simple Glasses: Exact Solutions in Infinite Dimensions (Cambridge University Press, 2020)
2020
Later among the works it cites.
K. Krishnamurthy, T. Can, and D. J. Schwab, arXiv preprint arXiv:2007.14823 (2020)
2020
Later among the works it cites.
S. Hwang and H. Ikeda, Physical Review E 101
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Li, S. Malladi, and S. Arora, arXiv preprint arXiv:2102.12470 (2021)
2021
Closest in time.
A. Bodin and N. Macris, Advances in Neural Information Processing Systems 34
2021
Closest in time.
F. Mignacco, P. Urbani, and L. Zdeborová, Machine Learning: Science and Technology 2
2021
Closest in time.
Y. Feng and Y. Tu, Proceedings of the National Academy of Sciences 118
2021
Closest in time.
R. Mandal and P. Sollich, Journal of Physics: Condensed Matter 33
2021
Closest in time.