Fetching the paper…
Reading the bibliography…
Optimizing deep neural networks (DNNs) often suffers from the ill-conditioned problem.
A simple weight decay can improve generalization
Anders Krogh and John A. Hertz · 1992
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
An Mei Chen, Haw minn Lu, and Robert Hecht-Nielsen · 1993
Earlier work this paper cites.
Effiicient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Riemannian geometry of Grassmann manifolds with a view on algorithmic computation
P.-A. Absil, R. Mahony, and R. Sepulchre · 2004
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
Joint diagonalization on the oblique manifold for independent component analysis
P. A. Absil and K. A. Gallivan · 2006
Earlier work this paper cites.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Maxout networks
Ian J. Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron C. Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Mean-normalized stochastic gradient for large-scale deep learning
Simon Wiesler, Alexander Richard, Ralf Schlüter, and Hermann Ney · 2014
Cited alongside, same era.
Scaling up natural gradient by sparsely factorizing the inverse fisher matrix
Roger B. Grosse and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Deeply-supervised nets
Chen-Yu Lee, Saining Xie, Patrick W. Gallagher, Zhengyou Zhang, and Zhuowen Tu · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan Salakhutdinov, and Nathan Srebro · 2015
Optimization on submanifolds of convolution kernels in cnns
Mete Ozay and Takayuki Okatani · 2016
Later among the works it cites.
Deconstructing the ladder network architecture
Mohammad Pezeshki, Linxi Fan, Philemon Brakel, Aaron C. Courville, and Yoshua Bengio · 2016
Later among the works it cites.
Improved techniques for training gans
Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Later among the works it cites.
Unsupervised and semi-supervised learning with categorical generative adversarial networks
Jost Tobias Springenberg · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Semi-supervised learning with ladder networks
Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, and Tapani Raiko · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martín Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Dizzyrnn: Reparameterizing recurrent neural networks for norm-preserving backpropagation
Victor Dorobantu, Per Andre Stromhaug, and Jess Renteria · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Later among the works it cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Later among the works it cites.
Riemannian approach to batch normalization
Minhyung Cho and Jaehyung Lee · 2017
Closest in time.
Generalized backpropagation, etude de cas: Orthogonality
Mehrtash Harandi and Basura Fernando · 2017
Closest in time.
Lei Huang, Xianglong Liu, Bo Lang, Admas Wei Yu, and Bo Li · 2017
Closest in time.
Triple generative adversarial nets
Chongxuan Li, Kun Xu, Jun Zhu, and Bo Zhang · 2017
Closest in time.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii · 2017
Closest in time.
Normalizing the normalizers: Comparing and extending network normalization schemes
Mengye Ren, Renjie Liao, Raquel Urtasun, Fabian H. Sinz, and Richard S. Zemel · 2017
Closest in time.
Neural machine translation with reconstruction
Zhaopeng Tu, Yang Liu, Lifeng Shang, Xiaohua Liu, and Hang Li · 2017
Closest in time.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Closest in time.
Normalized gradient with adaptive stepsize method for deep neural network training
Adams Wei Yu, Qihang Lin, Ruslan Salakhutdinov, and Jaime G. Carbonell · 2017
Closest in time.