Fetching the paper…
Reading the bibliography…
Orthogonal matrix has shown advantages in training Recurrent Neural Networks (RNNs), but such matrix is limited to be square for the hidden-to-hidden transformation in RNNs.
A simple weight decay can improve generalization
Anders Krogh and John A. Hertz · 1992
Earlier work this paper cites.
Effiicient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Mmse whitening and subspace whitening
Y. C. Eldar and A. V. Oppenheim · 2006
Earlier work this paper cites.
Special paraunitary matrices, cayley transform, and multidimensional orthogonal filter banks
Jianping Zhou, Minh N. Do, and Jelena Kovacevic · 2006
Earlier work this paper cites.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Notes on optimization on stiefel manifolds
Hemant D. Tagare · 2011
Earlier work this paper cites.
Projection-like retractions on matrix manifolds
Pierre-Antoine Absil and Jerome Malick · 2012
Earlier work this paper cites.
Orthogonalization of vectors with minimal adjustment
Paul H. Garthwaite, Frank Critchley, Karim Anaya-Izquierdo, and Emmanuel Mubwandarikwa · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Empirical arithmetic averaging over the compact stiefel manifold
Tetsuya Kaneko, Simone G. O. Fiori, and Toshihisa Tanaka · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
A feasible method for optimization with orthogonality constraints
Zaiwen Wen and Wotao Yin · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Linear dimensionality reduction: Survey, insights, and generalizations
John P. Cunningham and Zoubin Ghahramani · 2015
Cited alongside, same era.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, and Koray Kavukcuoglu · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Training deep networks with structured layers by matrix backpropagation
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Optimization on submanifolds of convolution kernels in cnns
Mete Ozay and Takayuki Okatani · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Later among the works it cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Catalin Ionescu, Orestis Vantzos, and Cristian Sminchisescu · 2015
Cited alongside, same era.
Semi-supervised learning with ladder networks
Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, and Tapani Raiko · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, Wei Liu, Yangqing Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martín Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Dizzyrnn: Reparameterizing recurrent neural networks for norm-preserving backpropagation
Victor Dorobantu, Per Andre Stromhaug, and Jess Renteria · 2016
Cited alongside, same era.
Later among the works it cites.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Closest in time.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Zakaria Mhammedi, Andrew D. Hellicar, Ashfaqur Rahman, and James Bailey · 2017
Closest in time.
Convolutional 2d lda for nonlinear dimensionality reduction
Feiping Nie Yuan Yuan Qi Wang, Zequn Qin · 2017
Closest in time.
Regularizing cnns with locally constrained decorrelations
Pau Rodríguez, Jordi Gonzàlez, Guillem Cucurull, Josep M. Gonfaus, and F. Xavier Roca · 2017
Closest in time.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Closest in time.
All you need is beyond a good init: Exploring better solution for training extremely deep convolutional neural networks with orthonormality and modulation
Di Xie, Jiang Xiong, and Shiliang Pu · 2017
Closest in time.
Block-normalized gradient method: An empirical study for training deep neural network
Adams Wei Yu, Lei Huang, Qihang Lin, Ruslan Salakhutdinov, and Jaime G. Carbonell · 2017
Closest in time.