Fetching the paper…
Reading the bibliography…
We introduce an efficient approach for optimization over orthogonal groups on highly parallel computation units such as GPUs or TPUs.
Unitary triangularization of a nonsymmetric matrix
A. S. Householder · 1958
Earlier work this paper cites.
Attractor dynamics and parallelism in a connectionist sequential machine
M. I. Jordan · 1990
Earlier work this paper cites.
Implementing automatic differentiation efficiently
D. Juedes and A. Griewank · 1990
Earlier work this paper cites.
Issues in parallel automatic differentiation
C. H. Bischof · 1991
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Numerical Linear Algebra
L. Trefethen and D. Bau · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
S. Hochreiter · 1998
Earlier work this paper cites.
Recognizing human actions: a local SVM approach
C. Schüldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Floating Point Operations in Matrix-vector Calculus
R. Hunger · 2005
Earlier work this paper cites.
Accumulating Householder transformations, revisited
T. Joffrain, T. M. Low, E. S. Quintana-Ortí, R. v. d. Geijn, and F. G. V. Zee · 2006
Earlier work this paper cites.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2007
Earlier work this paper cites.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, Second Edition
A. Griewank and A. Walther · 2008
Earlier work this paper cites.
Updating the qr factorization and the least squares problem
S. Hammarling and C. Lucas · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Mnist handwritten digit database. 2010
Y. LeCun, C. Cortes, and C. Burges · 2010
Earlier work this paper cites.
Geometric Methods and Applications: For Computer Science and Engineering
J. Gallier · 2011
Cited alongside, same era.
Notes on optimization on Stiefel manifolds
H. D. Tagare · 2011
Cited alongside, same era.
Stochastic gradient descent on Riemannian manifolds
S. Bonnabel · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Self-supervised visual planning with temporal skip connections
F. Ebert, C. Finn, A. X. Lee, and S. Levine · 2017
Later among the works it cites.
Efficient orthogonal parametrisation of recurrent neural networks using Householder reflections
Z. Mhammedi, A. Hellicar, A. Rahman, and J. Bailey · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Can we gain more from orthogonality regularizations in training deep networks?
N. Bansal, X. Chen, and Z. Wang · 2018
Later among the works it cites.
Orthogonal recurrent neural networks with scaled Cayley transform
K. Helfrich, D. Willmott, and Q. Ye · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Xingjian, Z. Chen, H. Wang, D.-Y. Yeung, W.-K. Wong, and W.-c. Woo · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Cited alongside, same era.
Dizzyrnn: Reparameterizing recurrent neural networks for norm-preserving backpropagation
V. Dorobantu, P. A. Stromhaug, and J. Renteria · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
M. Henaff, A. Szlam, and Y. LeCun · 2016
Cited alongside, same era.
Tunable efficient unitary neural networks (EUNN) and their application to RNN
L. Jing, Y. Shen, T. Dubcek, J. Peurifoy, S. A. Skirlo, M. Tegmark, and M. Soljacic · 2016
Cited alongside, same era.
Parallel matrix multiplication: A systematic journey
M. D. Schatz, R. A. Van de Geijn, and J. Poulson · 2016
Cited alongside, same era.
Orthogonal weight normalization: Solution to optimization over multiple dependent Stiefel manifolds in deep neural networks
L. Huang, X. Liu, B. Lang, A. W. Yu, Y. Wang, and B. Li · 2018
Later among the works it cites.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Later among the works it cites.
Sylvester normalizing flows for variational inference
R. Van Den Berg, L. Hasenclever, J. M. Tomczak, and M. Welling · 2018
Later among the works it cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
M. Artetxe and H. Schwenk · 2019
Later among the works it cites.
Computing the matrix exponential with an optimized taylor polynomial approximation
P. Bader, S. Blanes, and F. Casas · 2019
Later among the works it cites.
Trivializations for gradient-based optimization on manifolds
M. Lezcano Casado · 2019
Later among the works it cites.
Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group
M. Lezcano-Casado and D. Martínez-Rubio · 2019
Later among the works it cites.
Efficient Riemannian optimization on the Stiefel manifold via the Cayley transform
J. Li, F. Li, and S. Todorovic · 2020
Closest in time.
Parallel matrix computations, May 2020
M. Tuma · 2020
Closest in time.