Fetching the paper…
Reading the bibliography…
The goal of this tutorial is to introduce key models, algorithms, and open questions related to the use of optimization methods for solving problems arising in machine learning.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Towards an efficient sparsity exploiting newton method for minimization
Ph L Toint · 1981
Earlier work this paper cites.
The conjugate gradient method and trust regions in large scale optimization
Trond Steihaug · 1983
Earlier work this paper cites.
Neurocomputing: Foundations of research
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Predicting time series by a fully connected neural network trained by back propagation
C.R. Gent and C.P. Sheppard · 1992
Earlier work this paper cites.
Neural networks trained by analytically simulated damage states
M.F. Elkordy, K.C. Chang, and G.C. Lee · 1993
Earlier work this paper cites.
Representations of quasi N
R. H Byrd, J. Nocedal, and R. B Schnabel · 1994
Earlier work this paper cites.
Neural networks in civil engineering I: Principles and understanding
I. Flood and N. Kartam · 1994
Earlier work this paper cites.
Neural networks, a comprehensive foundation
Simon Haykin · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun and Yoshua Bengio · 1995
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V. N. Vapnik · 1995
Earlier work this paper cites.
Neural network design
Martin T Hagan, Howard B Demuth, Mark H Beale, and Orlando De Jesús · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Artificial neural networks for civil engineers: fundamentals and applications
N. Kartam, I. Flood, and J. H Garrett · 1997
Earlier work this paper cites.
Efficient BackProp
Y. LeCun, L. Bottou, G. Orr, and K. Muller · 1998
Earlier work this paper cites.
River stage forecasting using artificial neural networks
Konda Thirumalaiah and MC Deo · 1998
Earlier work this paper cites.
Solving the trust-region subproblem using the lanczos method
Nicholas IM Gould, Stefano Lucidi, Massimo Roma, and Philippe L Toint · 1999
Earlier work this paper cites.
Back-propagation neural network in tidal-level forecasting
Ching-Piao Tsai and Tsong-Lin Lee · 1999
Earlier work this paper cites.
Neural networks in civil engineering: 1989–2000
H. Adeli · 2001
Earlier work this paper cites.
Neural networks for wave forecasting
C. McDeo, A. Jha, A.S. Chaphekar, and K. Ravikant · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2003
Earlier work this paper cites.
Stochastic Programming
A. Ruszczynski and A. Shapiro, editors · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization
Yu. Nesterov · 2004
Earlier work this paper cites.
Introduction to Stochastic Search and Optimization: Estimation, Simulation, and Control
J.C. Spall · 2005
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
N. N. Schraudolph, J. Yu, and S. Gunter · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
O. Bousquet and L. Bottou · 2008
Earlier work this paper cites.
Approximation error bounds via rademacher’s complexity
G. Gnecco and M. Sanguineti · 2008
Cited alongside, same era.
Second-order stagewise backpropagation for hessian-matrix analyses and investigation of negative curvature
E. Mizutani and S. E. Dreyfus · 2008
Cited alongside, same era.
Fully connected network of superconducting qubits in a cavity
Dimitris I Tsomokos, Sahel Ashhab, and Franco Nori · 2008
Cited alongside, same era.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Cited alongside, same era.
Learning deep architectures for ai
Y. Bengio · 2009
Cited alongside, same era.
Model selection through sparse maximum likelihood estimation for multivariate gaussian or binary data
A. Bordes, L. Bottou, and P. Gallinari · 2009
Convergence of trust-region methods based on probabilistic models
Afonso S Bandeira, Katya Scheinberg, and Luis Nunes Vicente · 2014
Later among the works it cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Deep learning
Li Deng and Dong Yu · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reservoir computing approaches to recurrent neural network training
M. Lukoševičius and H. Jaeger · 2009
Cited alongside, same era.
Deep learning via hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Cited alongside, same era.
Deep learning and its applications to signal and information processing
Dong Yu and Li Deng · 2011
Cited alongside, same era.
Multi-column deep neural networks for image classification
Dan Ciregan, Ueli Meier, and Jürgen Schmidhuber · 2012
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data fitting
Michael P. Friedlander and Mark Schmidt · 2012
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Later among the works it cites.
On the saddle point problem for non-convex optimization
R. Pascanu, Y. N. Dauphin, S. Ganguli, and Y. Bengio · 2014
Later among the works it cites.
A Stochastic Quasi-Newton Method for Large-Scale Optimization
R.H. Byrd, S.L. Hansen, J. Nocedal, and Y.Singer · 2015
Later among the works it cites.
Global convergence rate analysis of unconstrained optimization methods based on probabilistic models
C. Cartis and K. Scheinberg · 2015
Later among the works it cites.
Stochastic optimization using a trust-region method and random models
C. Chen, M. Menickelly, and K. Scheinberg · 2015
Later among the works it cites.
Convergence rates of sub-sampled newton methods
M. A. Erdogdu and A. Montanari · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Zh. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Later among the works it cites.
A multi-batch l-bfgs method for machine learning
Albert S. Berahas, Jorge Nocedal, and Martin Takac · 2016
Later among the works it cites.
Convergence rate analysis of a stochastic trust region method for nonconvex optimization
J. Blanchet, C. Cartis, M. Menickelly, and K. Scheinberg · 2016
Later among the works it cites.
Exact and inexact subsampled newton methods for optimization
R. Bollapragada, R. Byrd, and J. Nocedal · 2016
Later among the works it cites.
Optimization Methods for Large-Scale Machine Learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
A self-correcting variable-metric algorithm for stochastic optimization
Frank Curtis · 2016
Later among the works it cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Later among the works it cites.
Stochastic block bfgs: Squeezing more curvature out of data
Robert Gower, Donald Goldfarb, and Peter Richtarik · 2016
Later among the works it cites.
Sub-sampled newton methods i: Globally convergent algorithms
F. Roosta-Khorasani and M. W. Mahoney · 2016
Later among the works it cites.
Sub-sampled newton methods ii: Local convergence rates
F. Roosta-Khorasani and M. W. Mahoney · 2016
Later among the works it cites.
Sub-sampled newton methods with non-uniform sampling
P. Xu, J. Yang, F. Roosta-Khorasani, Ch. Ré, and M. W. Mahoney · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
Complexity and global rates of trust-region methods based on probabilistic models
S. Gratton, C. W. Royer, L. N. Vicente, and Z. Zhang · 2017
Closest in time.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč · 2017
Closest in time.
Optimization Algorithms for Machine Learning Designed for Parallel and Distributed Environments
A. Yektamaram · 2017
Closest in time.