Fetching the paper…
Reading the bibliography…
The intuition that local flatness of the loss landscape is correlated with better generalization for deep neural networks (DNNs) has been explored for decades, spawning many different flatness measures.
Laws of information conservation (nongrowth) and aspects of the foundation of probability theory
L.A. Levin · 1974
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Keeping neural networks simple
Geoffrey E Hinton and Drew van Camp · 1993
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Radford M Neal · 1994
Earlier work this paper cites.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Some pac-bayesian theorems
David A McAllester · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Introduction to gaussian processes
David JC Mackay · 1998
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications
M. Li and P.M.B. Vitanyi · 2008
Earlier work this paper cites.
Expectation propagation of gaussian process classification and its application to gene expression analysis
Mingyue Tan · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
Kernels for vector-valued functions: A review
Mauricio A Alvarez, Lorenzo Rosasco, and Neil D Lawrence · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
The arrival of the frequent: how bias in genotype-phenotype maps can steer populations to local optima
Steffen Schaper and Ard A Louis · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The structure of the genotype–phenotype map strongly constrains the evolution of non-coding rna
Kamaludin Dingle, Steffen Schaper, and Ard A Louis · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Predicting the generalization gap in deep networks with margin distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2018
Later among the works it cites.
Input–output maps are strongly biased towards simple outputs
Kamaludin Dingle, Chico Q Camargo, and Ard A Louis · 2018
Later among the works it cites.
Random deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Toussi Kiani, and Seth Lloyd · 2018
Later among the works it cites.
How noise affects the hessian spectrum in overparameterized neural networks
Mingwei Wei and David J Schwab · 2019
Later among the works it cites.
A reparameterization-invariant flatness measure for deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Henning Petzka, Linara Adilova, Michael Kamp, and Cristian Sminchisescu · 2019
Later among the works it cites.
A scale invariant flatness measure for deep network minima
Akshay Rangamani, Nam H Nguyen, Abhishek Kumar, Dzung Phan, Sang H Chin, and Trac D Tran · 2019
Later among the works it cites.
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2019
Later among the works it cites.
Neural networks are a priori biased towards boolean functions with low entropy
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A Louis · 2019
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2019
Later among the works it cites.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
Modelling the influence of data structure on learning in neural networks
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Later among the works it cites.
Generalization bounds for deep learning
Guillermo Valle-Pérez and Ard A Louis · 2020
Later among the works it cites.
Understanding deep learning is also a job for physicists
Lenka Zdeborová · 2020
Later among the works it cites.
Generic predictions of output probability based on complexities of inputs and outputs
Kamaludin Dingle, Guillermo Valle Pérez, and Ard A Louis · 2020
Later among the works it cites.
Susanna Manrubia, José A Cuesta, Jacobo Aguirre, Sebastian E Ahnert, Lee Altenberg, Alejandro V Cano, Pablo Catalán, Ramon Diaz-Uriarte, Santiago F Elena, Juan Antonio García-Martín, et al · 2020
Later among the works it cites.
Is sgd a bayesian sampler? well, almost
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse, and Ard A Louis · 2021
Closest in time.
Learning by turning: Neural architecture aware optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, and Yisong Yue · 2021
Closest in time.