Fetching the paper…
Reading the bibliography…
Dropout regularization of deep neural networks has been a mysterious yet effective tool to prevent overfitting.
Scale Mixing of Symmetric Distributions with Zero Means
Beale, E. and Mallows, C · 1959
Earlier work this paper cites.
The M-Distribution: A General Formula of Intensity Distribution of Rapid Fading
Nakagami, M · 1960
Earlier work this paper cites.
Scale Mixtures of Normal Distributions
Andrews, D. F. and Mallows, C. L · 1974
Earlier work this paper cites.
Exponentially Decreasing Distributions for the Logarithm of Particle Size
Barndorff-Nielsen, O · 1977
Earlier work this paper cites.
Learning to Tell Two Spirals Apart
Lang, K. J. and Witbrock, M · 1988
Earlier work this paper cites.
Bayesian Variable Selection in Linear Regression
Mitchell, T. J. and Beauchamp, J. J · 1988
Earlier work this paper cites.
Bayesian Interpolation
MacKay, D. J · 1992
Earlier work this paper cites.
Variable Selection via Gibbs Sampling
George, E. I. and McCulloch, R. E · 1993
Earlier work this paper cites.
Bayesian Non-Linear Modeling for the Prediction Competition
MacKay, D. J · 1994
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1994
Earlier work this paper cites.
Variable Selection for Regression Models
Kuo, L. and Mallick, B · 1998
Earlier work this paper cites.
Bayesian Regression Analysis With Scale Mixtures of Normals
Steel, M. F · 2000
Earlier work this paper cites.
Sparse Bayesian Learning and the Relevance Vector Machine
Tipping, M. E · 2001
Earlier work this paper cites.
The Variational Bayesian EM Algorithm for Incomplete Data: with application to scoring graphical model structures
Beal, M. J. and Ghahramani, Z · 2003
Earlier work this paper cites.
Generalized Linear Models
Nelder, J. A. and Baker, R. J · 2004
Earlier work this paper cites.
Truncated Importance Sampling
Ionides, E. L · 2008
Earlier work this paper cites.
Handling Sparsity via the Horseshoe
Carvalho, C. M., Polson, N. G., and Scott, J. G · 2009
Earlier work this paper cites.
Improving Neural Networks by Preventing Co-Adaptation of Feature Detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Adaptive Dropout for Training Deep Neural Networks
Ba, J. and Frey, B · 2013
Earlier work this paper cites.
Understanding Dropout
Baldi, P. and Sadowski, P. J · 2013
Earlier work this paper cites.
Prediction of Breast Cancer Recurrence Using Classification Restricted Boltzmann Machine with Dropping
Tomczak, J. M · 2013
Earlier work this paper cites.
Dropout Training as Adaptive Regularization
Wager, S., Wang, S., and Liang, P. S · 2013
Earlier work this paper cites.
Regularization of Neural Networks Using DropConnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R · 2013
Earlier work this paper cites.
Fast Dropout Training
Wang, S. and Manning, C · 2013
Cited alongside, same era.
A Bayesian Encourages Dropout
Maeda, S.-i · 2014
Cited alongside, same era.
Analysis of Empirical MAP and Empirical Partially Bayes: Can They be Alternatives to Variational Bayes?
Nakajima, S. and Sugiyama, M · 2014
Cited alongside, same era.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Expectation Propagation as a Way of Life: A framework for Bayesian inference on partitioned data
Vehtari, A., Gelman, A., Sivula, T., Jylänki, P., Tran, D., Sahai, S., Blomstedt, P., Cunningham, J. P., Schiminovich, D., and Robert, C · 2014
Cited alongside, same era.
Altitude Training: Strong Bounds for Single-Layer Dropout
Wager, S., Fithian, W., Wang, S., and Liang, P. S · 2014
Swapout: Learning an Ensemble of Deep Architectures
Singh, S., Hoiem, D., and Forsyth, D · 2016
Later among the works it cites.
Generalized Dropout
Srinivas, S. and Venkatesh Babu, R · 2016
Later among the works it cites.
UCI Machine Learning Repository, 2017
Dheeru, D. and Karra Taniskidou, E · 2017
Later among the works it cites.
Bayesian Recurrent Neural Networks
Fortunato, M., Blundell, C., and Vinyals, O · 2017
Later among the works it cites.
Concrete Dropout
Gal, Y., Hron, J., and Kendall, A · 2017
Later among the works it cites.
Variational Inference for Sparse and Undirected Models
Ingraham, J. and Marks, D · 2017
Later among the works it cites.
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the Inductive Bias of Dropout
Helmbold, D. P. and Long, P. M · 2015
Cited alongside, same era.
Bayesian Dropout
Herlau, T., Mørup, M., and Schmidt, M. N · 2015
Cited alongside, same era.
Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
Hernández-Lobato, J. M. and Adams, R · 2015
Cited alongside, same era.
Variational Dropout and the Local Reparameterization Trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Cited alongside, same era.
Smart Regularization of Deep Architectures
Louizos, C · 2015
Cited alongside, same era.
A Statistical View of Deep Learning
Mohamed, S · 2015
Cited alongside, same era.
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N. R., Goyal, A., Bengio, Y., Courville, A., and Pal, C · 2017
Later among the works it cites.
Bayesian Compression for Deep Learning
Louizos, C., Ullrich, K., and Welling, M · 2017
Later among the works it cites.
Variational Dropout Sparsifies Deep Neural Networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Later among the works it cites.
Structured Bayesian Pruning via Log-Normal Multiplicative Noise
Neklyudov, K., Molchanov, D., Ashukha, A., and Vetrov, D. P · 2017
Later among the works it cites.
Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
Noh, H., You, T., Mun, J., and Han, B · 2017
Later among the works it cites.
Information Dropout: Learning Optimal Representations Through Noisy Computation
Achille, A. and Soatto, S · 2018
Closest in time.
Structured Variational Learning of Bayesian Neural Networks with Horseshoe Priors
Ghosh, S., Yao, J., and Doshi-Velez, F · 2018
Closest in time.
Variational Bayesian Dropout: Pitfalls and Fixes
Hron, J., Matthews, A., and Ghahramani, Z · 2018
Closest in time.
Learning Priors for Invariance
Nalisnick, E. and Smyth, P · 2018
Closest in time.
Posterior Concentration for Sparse Deep Learning
Polson, N. and Rockova, V · 2018
Closest in time.
Continuous Dropout
Shen, X., Tian, X., Liu, T., Xu, F., and Tao, D · 2018
Closest in time.
Variational Inference with Tail-adaptive f-Divergence
Wang, D., Liu, H., and Liu, Q · 2018
Closest in time.
Adaptive Dropout with Rademacher Complexity Regularization
Zhai, K. and Wang, H · 2018
Closest in time.
Fraternal Dropout
Zolna, K., Arpit, D., Suhubdy, D., and Bengio, Y · 2018
Closest in time.
β \beta -Dropout: a Unified Dropout
Liu, L., Luo, Y., Shen, X., Sun, M., and Li, B · 2019
Closest in time.
Deterministic Variational Inference for Robust Bayesian Neural Networks
Wu, A., Nowozin, S., Meeds, E., Turner, R. E., Hernández-Lobato, J. M., and Gaunt, A. L · 2019
Closest in time.