Fetching the paper…
Reading the bibliography…
Assisted by the availability of data and high performance computing, deep learning techniques have achieved breakthroughs and surpassed human performance empirically in difficult tasks, including object recognition, speech recognition, and natural language processing.
Y. L. Cun, O. Matan, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, L. D. Jacket, and H. S. Baird, “Handwritten zip code recognition with multilayer networks,” in [1990] Proceedings. 10th International Conference on Pattern Recognition , vol. ii, June 1990, pp. 35–40 vol.2
1990
Earlier work this paper cites.
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, Nov 1998
1998
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Ba and B. J. Frey, “Adaptive dropout for training deep neural networks,” in NIPS , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 , ser. NIPS’14, 2014, pp. 2933–2941
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, May 2015. [Online]. Available: http://dx.doi.org/10.1038/nature14539
2015
Earlier work this paper cites.
I. Goodfellow, O. Vinyals, and A. Saxe, “Qualitatively characterizing neural network optimization problems,” in International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
A. Choromanska, M. Henaff, M. Mathieu, G. Ben Arous, and Y. LeCun, “The loss surfaces of multilayer networks,” Journal of Machine Learning Research , vol. 38, pp. 192–204, 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” 2015
2015
Cited alongside, same era.
K. Kawaguchi, “Deep learning without poor local minima,” in Advances in Neural Information Processing Systems 29 , D. D. Lee, M. Sugiyama, U. V. Luxburg, I. Guyon, and R. Garnett, Eds., 2016, pp. 586–594
2016
Cited alongside, same era.
J. Manyika, M. Chui, M. Miremadi, J. Bughin, K. George, P. Willmott, and M. Dewhurst, “A future that works: Automation, employment, and productivity,” McKinsey Global Institute, Tech. Rep., 2017. [Online]. Available: https://www.mckinsey.com/~/media/McKinsey/Featured%20Insights/Digital%20Disruption/Harnessing%20automation%20for%20a%20future%20that%20works/MGI-A-future-that-works_Full-report.ashx
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . Cambridge, MA, USA: MIT Press, 2016, p. 276
2016
Cited alongside, same era.
M. Hardt, B. Recht, and Y. Singer, “Train faster, generalize better: Stability of stochastic gradient descent,” in Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 , ser. ICML’16, 2016, pp. 1225–1234
2016
Cited alongside, same era.
V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE , vol. 105, pp. 2295–2329, 2017
2017
Cited alongside, same era.
2018
Later among the works it cites.
L. P. K. Kenji Kawaguchi and Y. Bengio, “Generalization in deep learning,” in Mathematics of Deep Learning, Cambridge University Press, in preparation. Prepint avaliable as: MIT-CSAIL-TR-2018-014, Massachusetts Institute of Technology , 2018
2018
Later among the works it cites.
M. Olson, A. Wyner, and R. Berk, “Modern neural networks generalize on small data sets,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., 2018, pp. 3623–3632
2018
Later among the works it cites.
2018
Later among the works it cites.