Fetching the paper…
Reading the bibliography…
In many contexts, simpler models are preferable to more complex models and the control of this model complexity is the goal for many methods in machine learning such as regularization, hyperparameter tuning and architecture design.
Theory of uniform convergence of frequencies of events to their probabilities and problems of search for an optimal solution from empirical data(average risk minimization based on empirical data, showing relationship of problem to uniform convergence of averages toward expectation value)
AIA Chervonenkis and VN Vapnik · 1971
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Y. Chervonenkis · 1971
Earlier work this paper cites.
The method of ordered risk minimization, i
VN Vapnik and A Ya Chervonenkis · 1974
Earlier work this paper cites.
Learnability and the vapnik-chervonenkis dimension
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1989
Earlier work this paper cites.
Bayesian model comparison and backprop nets
David MacKay · 1991
Earlier work this paper cites.
Nonlinear total variation based noise removal algorithms
Leonid I Rudin, Stanley Osher, and Emad Fatemi · 1992
Earlier work this paper cites.
Autoencoders, minimum description length and helmholtz free energy
Geoffrey E Hinton and Richard Zemel · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
J. Schmidhuber · 1997
Earlier work this paper cites.
Vc dimension of neural networks
Eduardo D Sontag et al · 1998
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
V. Koltchinskii and D. Panchenko · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Vladimir Koltchinskii and Dmitriy Panchenko · 2000
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
Vladimir Koltchinskii · 2001
Earlier work this paper cites.
Model selection and error estimation
Peter L Bartlett, Stéphane Boucheron, and Gábor Lugosi · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Vladimir Koltchinskii and Dmitry Panchenko · 2002
Earlier work this paper cites.
Riemannian geometry and geometric analysis
Jürgen Jost and Jèurgen Jost · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Partial differential equations
Lawrence C Evans · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng · 2012
Earlier work this paper cites.
Rudin-osher-fatemi total variation denoising using split bregman
Pascal Getreuer · 2012
Earlier work this paper cites.
Dirichlet energy for analysis and synthesis of soft maps
Justin Solomon, Leonidas Guibas, and Adrian Butscher · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Degrees of freedom in deep neural networks
Tianxiang Gao and Vladimir Jojic · 2016
Earlier work this paper cites.
Deep Learning
Ian J. Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Geometric deep learning: going beyond euclidean data
Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Many paths to equilibrium: Gans do not need to decrease a divergence at every step
William Fedus, Mihaela Rosca, Balaji Lakshminarayanan, Andrew M Dai, Shakir Mohamed, and Ian Goodfellow · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
On convergence and stability of gans
Naveen Kodali, Jacob Abernethy, James Hays, and Zsolt Kira · 2017
Cited alongside, same era.
Why does deep and cheap learning work so well?
Henry W Lin, Max Tegmark, and David Rolnick · 2017
Cited alongside, same era.
The numerics of gans
Lars Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Cited alongside, same era.
Gradient descent gan optimization is locally stable
Vaishnavh Nagarajan and J Zico Kolter · 2017
Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Generalization in deep networks: The role of distance from initialization
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.
Adversarial robustness through local linearization
Chongli Qin, James Martens, Sven Gowal, Dilip Krishnan, Krishnamurthy Dvijotham, Alhussein Fawzi, Soham De, Robert Stanforth, and Pushmeet Kohli · 2019
Later among the works it cites.
Self-attention generative adversarial networks
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues · 2017
Cited alongside, same era.
Spectral norm regularization for improving the generalizability of deep learning
Yuichi Yoshida and Takeru Miyato · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 2018
Cited alongside, same era.
On gradient regularizers for mmd gans
Michael Arbel, Danica J Sutherland, Mikołaj Bińkowski, and Arthur Gretton · 2018
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Later among the works it cites.
In search of robust measures of generalization
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar, Ethan Caballero, Linbo Wang, Ioannis Mitliagkas, and Daniel M Roy · 2020
Later among the works it cites.
Robust learning with jacobian regularization
Judy Hoffman, Daniel A. Roberts, and Sho Yaida · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Later among the works it cites.
Implicit bias of gradient descent for mean squared error regression with wide neural networks
Hui Jin and Guido Montúfar · 2020
Later among the works it cites.
In Advances in Neural Information Processing Systems
Yoonho Lee, Juho Lee, Sung Ju Hwang, Eunho Yang, and Seungjin Choi · 2020
Later among the works it cites.
Neural complexity measures
Yoonho Lee, Juho Lee, Sung Ju Hwang, Eunho Yang, and Seungjin Choi · 2020
Later among the works it cites.
Chao Ma, Lei Wu, et al · 2020
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Later among the works it cites.
Training generative adversarial networks by solving ordinary differential equations
Chongli Qin, Yan Wu, Jost Tobias Springenberg, Andy Brock, Jeff Donahue, Timothy Lillicrap, and Pushmeet Kohli · 2020
Later among the works it cites.
A case for new neural network smoothness constraints
Mihaela Rosca, Theophane Weber, Arthur Gretton, and Shakir Mohamed · 2020
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Implicit gradient regularization
David G.T. Barrett and Benoit Dherin · 2021
Later among the works it cites.
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Mikhail Belkin · 2021
Later among the works it cites.
Towards deeper deep reinforcement learning
Johan Bjorck, Carla P Gomes, and Kilian Q Weinberger · 2021
Later among the works it cites.
A universal law of robustness via isoperimetry
Sébastien Bubeck and Mark Sellke · 2021
Later among the works it cites.
The geometric occam’s razor implicit in deep learning
Benoit Dherin, Michael Munn, and David G.T. Barrett · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phillip E Pope, Michael Moeller, and Tom Goldstein · 2021
Later among the works it cites.
Spectral normalisation for deep reinforcement learning: an optimisation perspective
Florin Gogianu, Tudor Berariu, Mihaela C Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Later among the works it cites.
The sobolev regularization effect of stochastic gradient descent
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.
Discretization drift in two-player games
Mihaela Rosca, Yan Wu, Benoit Dherin, and David G.T. Barrett · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David G.T. Barrett, and Soham De · 2021
Later among the works it cites.
Deep limits and a cut-off phenomenon for neural networks
Benny Avelin and Anders Karlsson · 2022
Closest in time.
Deep double descent via smooth interpolation
Matteo Gamba, Erik Englesson, Mårten Björkman, and Hossein Azizpour · 2022
Closest in time.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller, and Tom Goldstein · 2022
Closest in time.
Predicting generalization with degrees of freedom in neural networks
Erin Grant and Yan Wu · 2022
Closest in time.