Fetching the paper…
Reading the bibliography…
We show that a neural network with arbitrary depth and non-linearities, with dropout applied before every weight layer, is mathematically equivalent to an approximation to a well known Bayesian model.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Doubly stochastic variational Bayes for non-conjugate inference
Titsias, M. and Lázaro-Gredilla, M. (2014) · 1979
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Computing with infinite networks
Williams, C. K. (1997) · 1997
Earlier work this paper cites.
Marginalized kernels for biological sequences
Tsuda, K., Kin, T., and Asai, K. (2002) · 2002
Earlier work this paper cites.
Gaussian process dynamical models
Wang, J., Hertzmann, A., and Blei, D. M. (2005) · 2005
Earlier work this paper cites.
Neural probabilistic language models
Bengio, Y., Schwenk, H., Senécal, J.-S., Morin, F., and Gauvain, J.-L. (2006) · 2006
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I. (2006) · 2006
Earlier work this paper cites.
Sparse spectrum Gaussian process regression
Lázaro-Gredilla, M., Quiñonero-Candela, J., Rasmussen, C. E., and Figueiras-Vidal, A. R. (2010) · 2010
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Bayesian Gaussian process latent variable model
Titsias, M. and Lawrence, N. (2010) · 2010
Cited alongside, same era.
Variational Bayesian inference with stochastic search
Blei, D. M., Jordan, M. I., and Paisley, J. W. (2012) · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R. (2012) · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J. (2013) · 2013
Later among the works it cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Later among the works it cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Later among the works it cites.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. S. (2013) · 2013
Later among the works it cites.
Distributed variational inference in sparse Gaussian process regression and latent variable models
Gal, Y., van der Wilk, M., and Rasmussen, C. (2014) · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding dropout
Baldi, P. and Sadowski, P. J. (2013) · 2013
Cited alongside, same era.
Deep Gaussian processes
Damianou, A. and Lawrence, N. (2013) · 2013
Cited alongside, same era.
Gaussian processes for big data
Hensman, J., Fusi, N., and Lawrence, N. D. (2013) · 2013
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D. (2014) · 2014
Later among the works it cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2014) · 2014
Later among the works it cites.
Latent Gaussian processes for distribution estimation of multivariate categorical data
Gal, Y., Chen, Y., and Ghahramani, Z. (2015) · 2015
Closest in time.
Improving the Gaussian process sparse spectrum approximation by representing uncertainty in frequency inputs
Gal, Y. and Turner, R. (2015) · 2015
Closest in time.