Fetching the paper…
Reading the bibliography…
Learning a deep neural network requires solving a challenging optimization problem: it is a high-dimensional, non-convex and non-smooth minimization problem with a large number of terms.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David Rumelhart, Geoffrey Hinton, and Ronald Williams · 1986
Earlier work this paper cites.
The concave-convex procedure (CCCP)
Alan L. Yuille and Anand Rangarajan · 2002
Earlier work this paper cites.
Max-margin Markov networks
Benjamin Taskar, Carlos Guestrin, and Daphne Koller · 2003
Earlier work this paper cites.
Support vector machine learning for interdependent and structured output spaces
Ioannis Tsochantaridis, Thomas Hofmann, Thorsten Joachims, and Yasemin Altun · 2004
Earlier work this paper cites.
Optimal gradient-based learning using importance weights
Sepp Hochreiter and Klaus Obermayer · 2005
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Nicolas L Roux, Pierre-Antoine Manzagol, and Yoshua Bengio · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Training deep and recurrent networks with Hessian-free optimization
James Martens and Ilya Sutskever · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew Zeiler · 2012
Earlier work this paper cites.
Block-coordinate Frank-Wolfe optimization for structural SVMs
Simon Lacoste-Julien, Martin Jaggi, Mark Schmidt, and Patrick Pletscher · 2013
Earlier work this paper cites.
Riemannian metrics for neural networks
Yann Ollivier · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Duality between subgradient and conditional gradient methods
Francis Bach · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Natural neural networks
Guillaume Desjardins, Karen Simonyan, Razvan Pascanu, et al · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich, et al · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes · 2017
Later among the works it cites.
Reliably learning the ReLU in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2017
Later among the works it cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Later among the works it cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten · 2017
Later among the works it cites.
Learning to optimize
Ke Li and Jitendra Malik · 2017
Later among the works it cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Improper deep kernels
Uri Heinemann, Roi Livni, Elad Eban, Gal Elidan, and Amir Globerson · 2016
Cited alongside, same era.
A proximal method for composite minimization
Adrian S Lewis and Stephen J Wright · 2016
Cited alongside, same era.
Partial linearization based optimization for multi-class SVM
Pritish Mohapatra, Puneet Dokania, C. V. Jawahar, and M. Pawan Kumar · 2016
Cited alongside, same era.
Later among the works it cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2018
Closest in time.
Smooth loss functions for deep top-k classification
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2018
Closest in time.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Closest in time.
Proximal backpropagation
Thomas Frerix, Thomas Möllenhoff, Michael Moeller, and Daniel Cremers · 2018
Closest in time.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Closest in time.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
Faster convergence & generalization in DNNs
Gaurav Singh and John Shawe-Taylor · 2018
Closest in time.
WNGrad: Learn the learning rate in gradient descent
Xiaoxia Wu, Rachel Ward, and Léon Bottou · 2018
Closest in time.