Fetching the paper…
Reading the bibliography…
We examine Dropout through the perspective of interactions.
Generalized Additive Models
T. Hastie and R. Tibshirani · 1990
Earlier work this paper cites.
On overfitting and the effective number of hidden units
A. Weigend · 1994
Earlier work this paper cites.
Flexible smoothing with b-splines and penalties
P. H. Eilers and B. D. Marx · 1996
Earlier work this paper cites.
Early stopping-but when?
L. Prechelt · 1998
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
R. Caruana, S. Lawrence, and C. L. Giles · 2001
Earlier work this paper cites.
Simple incorporation of interactions into additive models
B. A. Coull, D. Ruppert, and M. Wand · 2001
Earlier work this paper cites.
Solving parity-n problems with feedforward neural networks
B. M. Wilamowski, D. Hunter, and A. Malinowski · 2003
Earlier work this paper cites.
An anova test for functional data
A. Cuevas, M. Febrero, and R. Fraiman · 2004
Earlier work this paper cites.
Diagnostics and Extrapolation in Machine Learning
G. Hooker · 2004
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Generalized functional anova diagnostics for high-dimensional functions of dependent variables
G. Hooker · 2007
Earlier work this paper cites.
Excitation dropout: Encouraging plasticity in deep neural networks
A. Zunino, S. A. Bargal, P. Morerio, J. Zhang, S. Sclaroff, and V. Murino · 2007
Earlier work this paper cites.
Sample sizes required to detect interactions between two binary fixed-effects in a mixed-effects linear regression model
A. C. Leon and M. Heo · 2009
Earlier work this paper cites.
Penalized wavelets: Embedding wavelets into semiparametric regression
M. Wand and J. T. Ormerod · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Earlier work this paper cites.
Intelligible models for classification and regression
Y. Lou, R. Caruana, and J. Gehrke · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
J. Ba and B. Frey · 2013
Earlier work this paper cites.
Understanding dropout
P. Baldi and P. J. Sadowski · 2013
Cited alongside, same era.
Accurate intelligible models with pairwise interactions
Y. Lou, R. Caruana, J. Gehrke, and G. Hooker · 2013
Cited alongside, same era.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. LeCun, and R. Fergus · 2013
Cited alongside, same era.
An empirical analysis of dropout in piecewise linear networks
D. Warde-Farley, I. J. Goodfellow, A. Courville, and Y. Bengio · 2013
Cited alongside, same era.
Avoiding pathologies in very deep networks
D. Duvenaud, O. Rippel, R. Adams, and Z. Ghahramani · 2014
Cited alongside, same era.
Deep neural networks are biased towards simple functions
G. De Palma, B. T. Kiani, and S. Lloyd · 2018
Later among the works it cites.
Statistical modeling, causal inference, and social science, Mar 2018
A. Gelman · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Later among the works it cites.
On the implicit bias of dropout
P. Mianjy, R. Arora, and R. Vidal · 2018
Later among the works it cites.
Dropout as a structured shrinkage prior
E. Nalisnick, J. M. Hernández-Lobato, and P. Smyth · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Tree boosting with xgboost-why does xgboost win" every" machine learning competition?
D. Nielsen · 2016
Cited alongside, same era.
Risk versus uncertainty in deep learning: Bayes, bootstrap and the dangers of dropout
I. Osband · 2016
Cited alongside, same era.
D. Selsam, M. Lamm, B. Bünz, P. Liang, L. de Moura, and D. L. Dill · 2018
Later among the works it cites.
Learning global additive explanations for neural nets using model distillation
S. Tan, R. Caruana, G. Hooker, P. Koch, and A. Gordo · 2018
Later among the works it cites.
Randomly projected additive gaussian processes for regression
I. A. Delbridge, D. S. Bindel, and A. G. Wilson · 2019
Later among the works it cites.
Visualizing the phate of neural networks
S. Gigante, A. S. Charles, S. Krishnaswamy, and G. Mishne · 2019
Later among the works it cites.
Sgd on neural networks learns functions of increasing complexity
P. Nakkiran, G. Kaplun, D. Kalimeris, T. Yang, B. L. Edelman, F. Zhang, and B. Barak · 2019
Later among the works it cites.
J. K. Tay and R. Tibshirani · 2019
Later among the works it cites.
On the structural sensitivity of deep convolutional networks to the directions of fourier basis functions
Y. Tsuzuku and I. Sato · 2019
Later among the works it cites.
A fourier perspective on model robustness in computer vision
D. Yin, R. G. Lopes, J. Shlens, E. D. Cubuk, and J. Gilmer · 2019
Later among the works it cites.
Multiplicative interactions and where to find them
S. M. Jayakumar, J. Menick, W. M. Czarnecki, J. Schwarz, J. Rae, S. Osindero, Y. W. Teh, T. Harley, and R. Pascanu · 2020
Closest in time.
Purifying interaction effects with the functional anova: An efficient algorithm for recovering identifiable additive models
B. Lengerich, S. Tan, C.-H. Chang, G. Hooker, and R. Caruana · 2020
Closest in time.
High-frequency component helps explain the generalization of convolutional neural networks
H. Wang, X. Wu, Z. Huang, and E. P. Xing · 2020
Closest in time.