Fetching the paper…
Reading the bibliography…
Strictly enforcing orthonormality constraints on parameter matrices has been shown advantageous in deep learning.
An introduction to differentiable manifolds and Riemannian geometry , volume 120
William M Boothby · 1986
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
SK Zavriev and FV Kostyuk · 1993
Earlier work this paper cites.
The geometry of algorithms with orthogonality constraints
Alan Edelman, Tomás A Arias, and Steven T Smith · 1998
Earlier work this paper cites.
Optimization algorithms exploiting unitary constraints
Jonathan H Manton · 2002
Earlier work this paper cites.
Learning algorithms utilizing quasi-geodesic flows on the stiefel manifold
Yasunori Nishimori and Shotaro Akaho · 2005
Earlier work this paper cites.
Special paraunitary matrices, Cayley transform, and multidi- mensional orthogonal filter banks
Jianping Zhou, Minh N. Do, and Jelena Kovacevic · 2006
Earlier work this paper cites.
Optimization algorithms on matrix manifolds
P-A Absil, Robert Mahony, and Rodolphe Sepulchre · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Projection-like retractions on matrix manifolds
P-A Absil and Jérôme Malick · 2012
Earlier work this paper cites.
Learning on the compact stiefel manifold by a cayley-transform-based pseudo-retraction map
Simone Fiori, Tetsuya Kaneko, and Toshihisa Tanaka · 2012
Earlier work this paper cites.
A feasible method for optimization with orthogonality constraints
Zaiwen Wen and Wotao Yin · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Reducing overfitting in deep networks by decorrelating representations
Michael Cogswell, Faruk Ahmed, Ross Girshick, Larry Zitnick, and Dhruv Batra · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Riemannian approach to batch normalization
Minhyung Cho and Jaehyung Lee · 2017
Later among the works it cites.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Later among the works it cites.
Orthogonal recurrent neural networks with scaled cayley transform
Kyle Helfrich, Devin Willmott, and Qiang Ye · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A framework of constraint preserving update schemes for optimization on stiefel manifold
Bo Jiang and Yu-Hong Dai · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Dizzyrnn: Reparameterizing recurrent neural networks for norm-preserving backpropagation
Victor Dorobantu, Per Andre Stromhaug, and Jess Renteria · 2016
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Cited alongside, same era.
Regularizing deep convolutional neural networks with a structured decorrelation constraint
Wei Xiong, Bo Du, Lefei Zhang, Ruimin Hu, and Dacheng Tao · 2016
Cited alongside, same era.
Orthogonal weight normalization: Solution to optimization over multiple dependent Stiefel manifolds in deep neural networks
Lei Huang, Xianglong Liu, Bo Lang, Adams Wei Yu, Yongliang Wang, and Bo Li
Cited in the paper.
Decorrelated batch normalization
Lei Huang, Dawei Yang, Bo Lang, and Jia Deng
Cited in the paper.
A riemannian conjugate gradient method for optimization on the stiefel manifold
Xiaojing Zhu · 2017
Later among the works it cites.
Can we gain more from orthogonality regularizations in training deep cnns?
Nitin Bansal, Xiaohan Chen, and Zhangyang Wang · 2018
Later among the works it cites.
High-order retractions on matrix manifolds using projected polynomials
Evan S Gawlik and Melvin Leok · 2018
Later among the works it cites.
Riemannian adaptive optimization methods
Gary Becigneul and Octavian-Eugen Ganea · 2019
Later among the works it cites.
Mario Lezcano-Casado and David Martínez-Rubio · 2019
Later among the works it cites.