Fetching the paper…
Reading the bibliography…
While dropout is known to be a successful regularization technique, insights into the mechanisms that lead to this success are still lacking.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
First-and second-order methods for learning: between steepest descent and newton’s method
Roberto Battiti · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
(not) bounding the true error
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
Pac-bayesian generalisation error bounds for gaussian process classification
Matthias Seeger · 2002
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2003
Earlier work this paper cites.
Generalized variance
S Kocherlakota and K Kocherlakota · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Pac-bayesian learning of linear classifiers
Pascal Germain, Alexandre Lacasse, François Laviolette, and Mario Marchand · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Pac-bayes bounds with data dependent priors
Emilio Parrado-Hernández, Amiran Ambroladze, John Shawe-Taylor, and Shiliang Sun · 2012
Earlier work this paper cites.
Second-order methods for neural networks: Fast and reliable training methods for multi-layer perceptrons
Adrian J Shepherd · 2012
Earlier work this paper cites.
Understanding dropout
Pierre Baldi and Peter J Sadowski · 2013
Earlier work this paper cites.
A pac-bayesian tutorial with a dropout bound
David McAllester · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Stefan Wager, Sida Wang, and Percy S Liang · 2013
Earlier work this paper cites.
Fast dropout training
Sida Wang and Christopher Manning · 2013
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
On the inductive bias of dropout
David P Helmbold and Philip M Long · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2015
Earlier work this paper cites.
On the properties of variational approximations of gibbs posteriors
Pierre Alquier, James Ridgway, and Nicolas Chopin · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Recurrent neural network for text classification with multi-task learning
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang · 2016
Cited alongside, same era.
Thuctc: an efficient chinese text classifier
Maosong Sun, Jingyang Li, Zhipeng Guo, Z Yu, Y Zheng, X Si, and Z Liu · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Dichotomize and generalize: Pac-bayesian binary activated deep neural networks
Gaël Letarte, Pascal Germain, Benjamin Guedj, and Francois Laviolette · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso A. Poggio, Alexander Rakhlin, and James Stokes · 2019
Later among the works it cites.
On dropout and nuclear norm regularization
Poorya Mianjy and Raman Arora · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Concrete dropout
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Cited alongside, same era.
Surprising properties of dropout in deep networks
David P Helmbold and Philip M Long · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Vaishnavh Nagarajan and J Zico Kolter · 2019
Later among the works it cites.
Dropout as a structured shrinkage prior
Eric Nalisnick, José Miguel Hernández-Lobato, and Padhraic Smyth · 2019
Later among the works it cites.
Information-theoretic generalization bounds for sgld via data-dependent estimates
Jeffrey Negrea, Mahdi Haghifam, Gintare Karolina Dziugaite, Ashish Khisti, and Daniel M Roy · 2019
Later among the works it cites.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
The intriguing role of module criticality in the generalization of deep networks
Niladri Chatterji, Behnam Neyshabur, and Hanie Sedghi · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin · 2020
Later among the works it cites.
Sharpened generalization bounds based on conditional mutual information and an application to noisy, iterative algorithms
Mahdi Haghifam, Jeffrey Negrea, Ashish Khisti, Daniel M Roy, and Gintare Karolina Dziugaite · 2020
Later among the works it cites.
Improving generalization by controlling label-noise information in neural network weights
Hrayr Harutyunyan, Kyle Reing, Greg Ver Steeg, and Aram Galstyan · 2020
Later among the works it cites.
Normalization techniques in training dnns: Methodology, analysis and application
Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao · 2020
Later among the works it cites.
Neurips 2020 competition: Predicting generalization in deep learning
Yiding Jiang, Pierre Foret, Scott Yak, M. Daniel Roy, Hossein Mobahi, Karolina Gintare Dziugaite, Samy Bengio, Suriya Gunasekar, Isabelle Guyon, and Behnam Neyshabur · 2020
Later among the works it cites.
How does weight correlation affect generalisation ability of deep neural networks?
Gaojie Jin, Xinping Yi, Liang Zhang, Lijun Zhang, Sven Schewe, and Xiaowei Huang · 2020
Later among the works it cites.
Ranking deep learning generalization using label variation in latent geometry graphs
Carlos Lassance, Louis Béthune, Myriam Bontonou, Mounia Hamidouche, and Vincent Gripon · 2020
Later among the works it cites.
Meta dropout: Learning to perturb latent features for generalization
Beom Hae Lee, Taewook Nam, Eunho Yang, and Ju Sung Hwang · 2020
Later among the works it cites.
On dropout, overfitting, and interaction effects in deep neural networks
Benjamin Lengerich, Eric P Xing, and Rich Caruana · 2020
Later among the works it cites.
Overparameterisation and worst-case generalisation: friend or foe?
Aditya Krishna Menon, Ankit Singh Rawat, and Sanjiv Kumar · 2020
Later among the works it cites.
On convergence and generalization of dropout training
Poorya Mianjy and Raman Arora · 2020
Later among the works it cites.
Representation based complexity measures for predicting generalization in deep learning
Parth Natekar and Manik Sharma · 2020
Later among the works it cites.
Pac-bayes analysis beyond the usual bounds
Omar Rivasplata, Ilja Kuzborskij, Csaba Szepesvári, and John Shawe-Taylor · 2020
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Later among the works it cites.
On the generalization effects of linear transformations in data augmentation
Sen Wu, Hongyang Zhang, Gregory Valiant, and Christopher Ré · 2020
Later among the works it cites.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Later among the works it cites.
Contextual dropout: An efficient sample-dependent dropout module
Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, and Mingyuan Zhou · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Rethinking the pruning criteria for convolutional neural network
Zhongzhan Huang, Wenqi Shao, Xinjiang Wang, Liang Lin, and Ping Luo · 2021
Later among the works it cites.
Tighter risk certificates for neural networks
Marıa Pérez-Ortiz, Omar Rivasplata, John Shawe-Taylor, and Csaba Szepesvári · 2021
Later among the works it cites.
Autodropout: Learning dropout patterns to regularize deep networks
Hieu Pham and Quoc Le · 2021
Later among the works it cites.
Towards understanding and improving dropout in game theory
Hao Zhang, Sen Li, YinChao Ma, Mingjie Li, Yichen Xie, and Quanshi Zhang · 2021
Later among the works it cites.
Enhancing adversarial training with second-order statistics of weights
Gaojie Jin, Xinping Yi, Wei Huang, Sven Schewe, and Xiaowei Huang · 2022
Closest in time.