Fetching the paper…
Reading the bibliography…
We introduce a novel approach to perform first-order optimization with orthogonal and unitary constraints.
Curvatures of left invariant metrics on lie groups
Milnor, J · 1976
Earlier work this paper cites.
Riemannian Geometry
do Carmo, M · 1992
Earlier work this paper cites.
Timit acoustic-phonetic continuous speech corpus
S Garofolo, J., Lamel, L., M Fisher, W., Fiscus, J., S. Pallett, D., L. Dahlgren, N., and Zue, V · 1992
Earlier work this paper cites.
Riemannian svrg: Fast stochastic optimization on riemannian manifolds
Zhang, H., Reddi, S. J., and Sra, S · 1992
Earlier work this paper cites.
Geometric Optimization Methods for Adaptive Filtering
Smith, S. T · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The geometry of algorithms with orthogonality constraints
Edelman, A., Arias, T. A., and Smith, S. T · 1998
Earlier work this paper cites.
Optimization algorithms exploiting unitary constraints
Manton, J. H · 2002
Earlier work this paper cites.
Nineteen dubious ways to compute the exponential of a matrix, twenty-five years later
Moler, C. and Van Loan, C · 2003
Earlier work this paper cites.
Steepest descent algorithms for optimization under unitary matrix constraint
Abrudan, T. E., Eriksson, J., and Koivunen, V · 2008
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
Absil, P.-A., Mahony, R., and Sepulchre, R · 2009
Earlier work this paper cites.
The scaling and squaring method for the matrix exponential revisited
Higham, N. J · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Stochastic gradient descent on riemannian manifolds
Bonnabel, S · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
Lie Groups, Lie Algebras, and Representations: An Elementary Introduction
Hall, B · 2015
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
Wisdom, S., Powers, T., Hershey, J., Le Roux, J., and Atlas, L · 2016
Later among the works it cites.
Huang, L., Liu, X., Lang, B., Yu, A. W., Wang, Y., and Li, B · 2017
Later among the works it cites.
Learning unitary operators with help from u (n)
Hyland, S. L. and Rätsch, G · 2017
Later among the works it cites.
Tunable efficient unitary neural networks (eunn) and their application to rnns
Jing, L., Shen, Y., Dubcek, T., Peurifoy, J., Skirlo, S., LeCun, Y., Tegmark, M., and Soljačić, M · 2017
Later among the works it cites.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Mhammedi, Z., Hellicar, A., Rahman, A., and Bailey, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A framework of constraint preserving update schemes for optimization on stiefel manifold
Jiang, B. and Dai, Y.-H · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Le, Q. V., Jaitly, N., and Hinton, G. E · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Cited alongside, same era.
Global rates of convergence for nonconvex optimization on manifolds
Boumal, N., Absil, P.-A., and Cartis, C · 2016
Cited alongside, same era.
Generalized backpropagation, étude de cas: Orthogonality
Harandi, M. and Fernando, B · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Henaff, M., Szlam, A., and LeCun, Y · 2016
Cited alongside, same era.
Sato, H., Kasai, H., and Mishra, B · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Vorontsov, E., Trabelsi, C., Kadoury, S., and Pal, C · 2017
Later among the works it cites.
Can we gain more from orthogonality regularizations in training deep cnns?
Bansal, N., Chen, X., and Wang, Z · 2018
Later among the works it cites.
Orthogonal recurrent neural networks with scaled Cayley transform
Helfrich, K., Willmott, D., and Ye, Q · 2018
Later among the works it cites.
Geometry aware constrained optimization techniques for deep learning
Kumar Roy, S., Mhammedi, Z., and Harandi, M · 2018
Later among the works it cites.
Complex unitary recurrent neural networks using scaled cayley transform
Maduranga, K. D. G., Helfrich, K., and Ye, Q · 2018
Later among the works it cites.
Riemannian adaptive optimization methods
Becigneul, G. and Ganea, O.-E · 2019
Closest in time.