Fetching the paper…
Reading the bibliography…
In this paper, we show that feedforward and recurrent neural networks exhibit an outer product derivative structure but that convolutional neural networks do not.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” nature
1986
Earlier work this paper cites.
C. Bishop, “Exact calculation of the hessian matrix for the multilayer perceptron,” Neural Computation
1992
Earlier work this paper cites.
H. Drucker and Y. Le Cun, “Improving generalization performance using double backpropagation,” IEEE Transactions on Neural Networks
1992
Earlier work this paper cites.
J. E. Moody, “The effective number of parameters: An analysis of generalization and regularization in nonlinear learning systems,” in Advances in neural information processing systems
1992
Earlier work this paper cites.
W. L. Buntine and A. S. Weigend, “Computing second derivatives in feed-forward networks: A review,” IEEE transactions on Neural Networks
1994
Earlier work this paper cites.
B. A. Pearlmutter, “Fast exact multiplication by the hessian,” Neural computation
1994
Earlier work this paper cites.
World Scientific, 2007
V. G. Ivancevic and T. T. Ivancevic, Applied differential geometry: a modern introduction · 2007
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy layer-wise training of deep networks,” in Advances in neural information processing systems
2007
Earlier work this paper cites.
N. N. Schraudolph, J. Yu, and S. Günter, “A stochastic quasi-Newton method for online convex optimization,” in Artificial Intelligence and Statistics
2007
Earlier work this paper cites.
E. Mizutani and S. E. Dreyfus, “Second-order stagewise backpropagation for hessian-matrix analyses and investigation of negative curvature,” Neural Networks
2008
Earlier work this paper cites.
J. Martens, “Deep learning via hessian-free optimization,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10)
2010
Earlier work this paper cites.
R. H. Byrd, G. M. Chin, W. Neveitt, and J. Nocedal, “On the use of stochastic Hessian information in optimization methods for machine learning,” SIAM Journal on Optimization
2011
Cited alongside, same era.
J. Martens, I. Sutskever, and K. Swersky, “Estimating the Hessian by back-propagating curvature,” in Proceedings of the 29th International Coference on International Conference on Machine Learning
2012
Cited alongside, same era.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International conference on machine learning
2013
Cited alongside, same era.
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio, “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Advances in neural information processing systems
2014
Cited alongside, same era.
MIT press Cambridge, 2016
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning · 2016
Later among the works it cites.
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer, “A stochastic quasi-Newton method for large-scale optimization,” SIAM Journal on Optimization
2016
Later among the works it cites.
P. Moritz, R. Nishihara, and M. Jordan, “A linearly-convergent stochastic L-BFGS algorithm,” in Artificial Intelligence and Statistics
2016
Later among the works it cites.
N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami, “Distillation as a defense to adversarial perturbations against deep neural networks,” in 2016 IEEE Symposium on Security and Privacy (SP)
2016
Later among the works it cites.
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?,” in Advances in neural information processing systems
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980
2014
Cited alongside, same era.
A. Mokhtari and A. Ribeiro, “Res: Regularized stochastic BFGS algorithm,” IEEE Transactions on Signal Processing
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Cited alongside, same era.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Closest in time.