Fetching the paper…
Reading the bibliography…
The largely successful method of training neural networks is to learn their weights using some variant of stochastic gradient descent (SGD).
On analytical methods in probability theory
Andrey Nikolaevich Kolmogorov · 1931
Earlier work this paper cites.
On stochastic differential equations
Kiyosi Itô · 1951
Earlier work this paper cites.
Using fast weights to deblur old memories
Geoffrey E Hinton and David C Plaut · 1987
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
David J. C. MacKay · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: an alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
The variational formulation of the Fokker–Planck equation
Richard Jordan, David Kinderlehrer, and Felix Otto · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Statistical mechanical approaches to models with many poorly known parameters
Kevin S. Brown and James P. Sethna · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Sloppy-model universality class and the Vandermonde matrix
Joshua J. Waterfall, Fergal P. Casey, Ryan N. Gutenkunst, Kevin S. Brown, Christopher R. Myers, Piet W. Brouwer, Veit Elser, and James P. Sethna · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Delving deep into rectifiers: surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Why M heads are better than one: training a diverse ensemble of deep networks
Stefan Lee, Senthil Purushwalkam, Michael Cogswell, David Crandall, and Dhruv Batra · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li · 2015
Earlier work this paper cites.
LSUN: construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Deep learning with elastic averaging SGD
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Cited alongside, same era.
Dropout as a Bayesian approximation: representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Second-order optimization for neural networks
James Martens · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Fast context adaptation via meta-learning
Luisa Zintgraf, Kyriacos Shiarli, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2016
Cited alongside, same era.
Continual reinforcement learning with complex synapses
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2018
Later among the works it cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Later among the works it cites.
Film: visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Later among the works it cites.
Empirical analysis of the Hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Terrance DeVries and Graham W. Taylor · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Google vizier: a service for black-box optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D. Sculley · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Hypernetworks
David Ha, Andrew M. Dai, and Quoc V. Le · 2017
Cited alongside, same era.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2018
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Later among the works it cites.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Later among the works it cites.
Synaptic weight decay with selective consolidation enables fast learning without catastrophic forgetting
Pascal Leimer, Michael Herzog, and Walter Senn · 2019
Later among the works it cites.
Deep learning theory review: An optimal control and dynamical systems perspective
Guan-Horng Liu and Evangelos A Theodorou · 2019
Later among the works it cites.
A simple baseline for Bayesian uncertainty in deep learning
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
K for the price of 1: parameter-efficient multi-task and transfer learning
Pramod Kaushik Mudrakarta, Mark Sandler, Andrey Zhmoginov, and Andrew Howard · 2019
Later among the works it cites.
Learning implicitly recurrent CNNs through parameter sharing
Pedro Savarese and Michael Maire · 2019
Later among the works it cites.
Shakedrop regularization for deep residual learning
Yoshihiro Yamada, Masakazu Iwamura, Takuya Akiba, and Koichi Kise · 2019
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
Carlo Baldassi, Fabrizio Pittorino, and Riccardo Zecchina · 2020
Closest in time.
Meta-learning with warped gradient descent
Sebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin, Hujun Yin, and Raia Hadsell · 2020
Closest in time.
Deep ensembles: a loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2020
Closest in time.
Training batchnorm and only batchnorm: on the expressive power of random features in CNNs
Jonathan Frankle, David J. Schwab, and Ari S. Morcos · 2020
Closest in time.
Multiplicative interactions and where to find them
Siddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osindero, Yee Whye Teh, Tim Harley, and Razvan Pascanu · 2020
Closest in time.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Closest in time.
Entropic gradient descent algorithms and wide flat minima
Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Enrico M. Malatesta, Gabriele Perugini, Carlo Baldassi, Matteo Negri, Elizaveta Demyanenko, and Riccardo Zecchina · 2020
Closest in time.
Continual learning with hypernetworks
Johannes von Oswald, Christian Henning, João Sacramento, and Benjamin F. Grewe · 2020
Closest in time.
BatchEnsemble: an alternative approach to efficient ensemble and lifelong learning
Yeming Wen, Dustin Tran, and Jimmy Ba · 2020
Closest in time.