Fetching the paper…
Reading the bibliography…
Recent works have highlighted scale invariance or symmetry that is present in the weight space of a typical deep network and the adverse effect that it has on the Euclidean gradient based stochastic gradient descent optimization.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
The geometry of algorithms with orthogonality constraints
Edelman, A., Arias, T.A., and Smith, S.T · 1998
Earlier work this paper cites.
Optimization algorithms exploiting unitary constraints
Manton, J.H · 2002
Earlier work this paper cites.
Riemannian geometry of Grassmann manifolds with a view on algorithmic computation
Absil, P.-A., Mahony, R., and Sepulchre, R · 2004
Earlier work this paper cites.
Optimization Algorithms on Matrix Manifolds
Absil, P.-A., Mahony, R., and Sepulchre, R · 2008
Earlier work this paper cites.
Lecture notes
Hinton, G · 2008
Earlier work this paper cites.
Semantic object classes in video: A high-definition ground truth database
Brostow, G., Fauqueur, J., and Cipolla, R · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Cited alongside, same era.
Low-rank optimization on the cone of positive semidefinite matrices
Journée, M., Bach, F., Absil, P.-A., and Sepulchre, R · 2010
Cited alongside, same era.
Stochastic gradient descent on Riemannian manifolds
Bonnabel, S · 2013
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Cited alongside, same era.
Mishra, B. and Sepulchre, R · 2014
Later among the works it cites.
Low-rank matrix completion via preconditioned optimization on the Grassmann manifold
Boumal, N. and Absil, P.-A · 2015
Closest in time.
Desjardins, G., Simonyan, K., Pascanu, R., and Kavukcuoglu, K · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Closest in time.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Closest in time.
Path-sgd: Path-normalized optimization in deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Manopt, a matlab toolbox for optimization on manifolds
Boumal, N., Mishra, B., Absil, P.-A., and Sepulchre, R · 2014
Cited alongside, same era.
SegNet: a deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling
Badrinarayanan, V., Handa, A., and Cipolla, R
Cited in the paper.
Understanding symmetries in deep networks
Badrinarayanan, V., Mishra, B., and Cipolla, R
Cited in the paper.
Riemannian metrics for neural networks I: Feedforward networks
Ollivier, Y
Cited in the paper.
Riemannian metrics for neural networks II: Recurrent networks and learning symbolic data sequences
Ollivier, Y
Cited in the paper.
Neyshabur, B., Salakhutdinov, R., and Srebro, N · 2015
Closest in time.