Fetching the paper…
Reading the bibliography…
Convolutional neural networks (CNNs) are trained using stochastic gradient descent (SGD)-based optimizers.
H. Rosenbrock, “An automatic method for finding the greatest or least value of a function,” The Computer Journal , vol. 3, no. 3, pp. 175–184, 1960
1960
Earlier work this paper cites.
L. Bottou, “Stochastic gradient learning in neural networks,” Proceedings of Neuro-Nımes , vol. 91, no. 8, p. 12, 1991
1991
Earlier work this paper cites.
A. L. Blum and R. L. Rivest, “Training a 3-node neural network is np-complete,” Neural Networks , vol. 5, no. 1, pp. 117–127, 1992
1992
Earlier work this paper cites.
G. E. Moore, “Cramming more components onto integrated circuits,” Proceedings of the IEEE , vol. 86, no. 1, pp. 82–85, 1998
1998
Earlier work this paper cites.
N. Qian, “On the momentum term in gradient descent learning algorithms,” Neural networks , vol. 12, no. 1, pp. 145–151, 1999
1999
Earlier work this paper cites.
B. C. Csáji et al. , “Approximation with artificial neural networks,” Faculty of Sciences, Etvs Lornd University, Hungary , vol. 24, no. 48, p. 7, 2001
2001
Earlier work this paper cites.
C. M. Bishop, Pattern recognition and machine learning . springer, 2006
2006
Earlier work this paper cites.
S. Gopisetty, S. Agarwala, E. Butler, D. Jadav, S. Jaquet, M. Korupolu, R. Routray, P. Sarkar, A. Singh, M. Sivan-Zimet et al. , “Evolution of storage management: Transforming raw data into information,” IBM Journal of Research and Development , vol. 52, no. 4.5, pp. 341–352, 2008
2008
Earlier work this paper cites.
G. E. Hinton, “Deep belief networks,” Scholarpedia , vol. 4, no. 5, p. 5947, 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Department of Computer Science, University of Toronto , 2009
2009
Earlier work this paper cites.
D. Yu and L. Deng, “Deep learning and its applications to signal and information processing [exploratory dsp],” IEEE Signal Processing Magazine , vol. 28, no. 1, pp. 145–154, 2010
2010
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization.” Journal of machine learning research , vol. 12, no. 7, 2011
2011
Earlier work this paper cites.
A. Khosla, N. Jayadevaprakash, B. Yao, and F.-F. Li, “Novel dataset for fgvc: Stanford dogs,” in San Diego: CVPR Workshop on FGVC , vol. 1, no. 2, 2011, p.
2011
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” California Institute of Technology, Tech. Rep. CNS-TR-2011-001, 2011
2011
Earlier work this paper cites.
D. C. Cireşan, U. Meier, L. M. Gambardella, and J. Schmidhuber, “Deep big multilayer perceptrons for digit recognition,” in Neural networks: tricks of the trade . Springer, 2012, pp. 581–598
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems , F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, Eds., vol. 25. Curran Associates, Inc., 2012, pp. 1097–1105
2012
Earlier work this paper cites.
G. Hinton, N. Srivastava, and K. Swersky, “Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,” Cited on , vol. 14, no. 8, 2012
2012
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International conference on machine learning , 2013, pp. 1310–1318
2013
Earlier work this paper cites.
C. Szegedy, A. Toshev, and D. Erhan, “Deep neural networks for object detection,” in Advances in neural information processing systems , 2013, pp. 2553–2561
2013
Earlier work this paper cites.
N. Wang and D.-Y. Yeung, “Learning a deep compact image representation for visual tracking,” in Advances in neural information processing systems , 2013, pp. 809–817
2013
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International conference on machine learning . PMLR, 2013, pp. 1139–1147
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in Proceedings of the IEEE international conference on computer vision workshops , 2013, pp. 554–561
2013
Cited alongside, same era.
2013
Cited alongside, same era.
C. P. Chen and C.-Y. Zhang, “Data-intensive applications, challenges, techniques and technologies: A survey on big data,” Information sciences , vol. 275, pp. 314–347, 2014
2014
Cited alongside, same era.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Advances in neural information processing systems , vol. 3, no. 06, 2014
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” in International conference on learning representations , 2017, p.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Cited alongside, same era.
J. Schmidhuber, “Deep learning in neural networks: An overview,” Neural networks , vol. 61, pp. 85–117, 2015
2015
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016, http://www.deeplearningbook.org
2016
Cited alongside, same era.
J. Zabalza, J. Ren, J. Zheng, H. Zhao, C. Qing, Z. Yang, P. Du, and S. Marshall, “Novel segmented stacked autoencoder for effective dimensionality reduction and feature extraction in hyperspectral imaging,” Neurocomputing , vol. 185, pp. 1–10, 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
2018
Later among the works it cites.
G. Wang, W. Li, M. A. Zuluaga, R. Pratt, P. A. Patel, M. Aertsen, T. Doel, A. L. David, J. Deprest, S. Ourselin et al. , “Interactive medical image segmentation using deep learning with image-specific fine tuning,” IEEE transactions on medical imaging , vol. 37, no. 7, pp. 1562–1573, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
F. Yu, D. Wang, E. Shelhamer, and T. Darrell, “Deep layer aggregation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 2403–2412
2018
Later among the works it cites.
Z. Zhang and H. Peng, “Deeper and wider siamese networks for real-time visual tracking,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4591–4600
2019
Later among the works it cites.
M. Paoletti, J. Haut, J. Plaza, and A. Plaza, “Deep learning classifiers for hyperspectral imaging: A review,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 158, pp. 279–317, 2019
2019
Later among the works it cites.
Y. N. Dauphin and S. Schoenholz, “Metainit: Initializing learning by learning to initialize,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Iscen, G. Tolias, Y. Avrithis, and O. Chum, “Label propagation for deep semi-supervised learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5070–5079
2019
Later among the works it cites.
A. Mirzaei, V. Pourahmadi, M. Soltani, and H. Sheikhzadeh, “Deep feature selection using a teacher-student network,” Neurocomputing , vol. 383, pp. 396–408, 2020
2020
Later among the works it cites.
D.-X. Zhou, “Universality of deep convolutional neural networks,” Applied and computational harmonic analysis , vol. 48, no. 2, pp. 787–794, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Dubey, S. Chakraborty, S. Roy, S. Mukherjee, S. Singh, and B. Chaudhuri, “diffgrad: An optimization method for convolutional neural networks.” IEEE transactions on neural networks and learning systems , vol. 31, no. 11, pp. 4500–4511, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han, “On the variance of the adaptive learning rate and beyond,” in Proceedings of the Eighth International Conference on Learning Representations (ICLR 2020) , April 2020, p.
2020
Later among the works it cites.
J. Zhuang, T. Tang, Y. Ding, S. Tatikonda, N. Dvornek, X. Papademetris, and J. Duncan, “Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,” Conference on Neural Information Processing Systems , 2020
2020
Later among the works it cites.