J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research , vol. 12, no. Jul, pp. 2121–2159, 2011
2011
Cited alongside, same era.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in ICML Workshop on Deep Learning for Audio, Speech and Language Processing , 2013
2013
Cited alongside, same era.
B. Neyshabur, R. R. Salakhutdinov, and N. Srebro, “Path-SGD: Path-normalized optimization in deep neural networks,” in Advances in Neural Information Processing Systems , 2015, pp. 2422–2430
2015
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Cited alongside, same era.
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro, “Exploring generalization in deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 5947–5956
2017
Cited alongside, same era.
M. Unser, J. Fageot, and J. P. Ward, “Splines are universal solutions of linear inverse problems with generalized TV regularization,” SIAM Review , vol. 59, no. 4, pp. 769–793, 2017
2017
Cited alongside, same era.
J. M. Klusowski and A. R. Barron, “Approximation by combinations of ReLU and squared ReLU ridge functions with ℓ 1 \ell^{1} and ℓ 0 \ell^{0} controls,” IEEE Transactions on Information Theory , vol. 64, no. 12, pp. 7649–7656, 2018
2018
Cited alongside, same era.