Fetching the paper…
Reading the bibliography…
Bayesian deep learning offers a principled way to address many issues concerning safety of artificial intelligence (AI), such as model uncertainty,model interpretability, and prediction bias.
1902
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
Metropolis, N., Rosenbluth, A., Rosenbluth, M., Teller, A., and Teller, E. (1953), “Equation of state calculations by fast computing machines,” Journal of Chemical Physics
1953
Earlier work this paper cites.
Sutton, R. S. (1986), “Two Problems with Backpropagation and Other Steepest-Descent Learning Procedures for Networks,” in Proceedings of the Eighth Annual Conference of the Cognitive Science Society
1986
Earlier work this paper cites.
Qian, N. (1999), “On the momentum term in gradient descent learning algorithms,” Neural Networks
1999
Earlier work this paper cites.
Haaro, H., Saksman, E., and Tamminen, J. (2001), “An Adaptive Metropolis Algorithm,” Bernoulli
2001
Earlier work this paper cites.
Mattingly, J., Stuartb, A., and Highamc, D. (2002), “Ergodicity for SDEs and Approximations: Locally Lipschitz Vector Fields and Degenerate Noise,” Stochastic Processes and their Applications
2002
Earlier work this paper cites.
Song, Q., Sun, Y., Ye, M., and Liang, F. (2020), “Extended Stochastic Gradient MCMC for Large-Scale Bayesian Variable Selection,” arXiv:
2002
Earlier work this paper cites.
Duchi, J., Hazan, E., and Singer, Y. (2011), “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research
2011
Earlier work this paper cites.
Girolami, M. and Calderhead, B. (2011), “Riemann manifold Langevin and Hamiltonian Monte Carlo methods (with discussion),” Journal of the Royal Statistical Society, Series B
2011
Earlier work this paper cites.
Welling, M. and Teh, Y. W. (2011), “Bayesian Learning via Stochastic Gradient Langevin Dynamics,” in ICML
2011
Earlier work this paper cites.
Ahn, S., Korattikara, A., and Welling, M. (2012), “Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring,” in ICML
2012
Earlier work this paper cites.
Tieleman, T. and Hinton, G. (2012), “Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Machine Learning
2012
Cited alongside, same era.
Zeiler, M. D. (2012), “ADADELTA: An Adaptive Learning Rate Method,” CoRR
2012
Cited alongside, same era.
Patterson, S. and Teh, Y. W. (2013), “Stochastic Gradient Riemannian Langevin Dynamics on the Probability Simplex,” in Advances in Neural Information Processing Systems 26
2013
Cited alongside, same era.
Chen, T., Fox, E. B., and Guestrin, C. (2014), “Stochastic Gradient Hamiltonian Monte Carlo,” in ICML
2014
Cited alongside, same era.
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014), “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization,” in Advances in Neural Information Processing Systems 27
Vollmer, S. J., Zygalakis, K. C., and Teh, Y. W. (2016), “Exploration of the (Non-)Asymptotic Bias and Variance of Stochastic Gradient Langevin Dynamics,” Journal of Machine Learning Research
2016
Later among the works it cites.
2017
Later among the works it cites.
Kendall, A. and Gal, Y. (2017), “What uncertainties do we need in Bayesian deep learning for computer vision,” in The 31st Conference on Neural Information Processing Systems (NIPS 2017)
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
Kingma, D. and Ba, J. (2014), “Adam: A Method for Stochastic Optimization,” International Conference on Learning Representations
2014
Cited alongside, same era.
Sato, I. and Nakagawa, H. (2014), “Approximation Analysis of Stochastic Gradient Langevin Dynamics by using Fokker-Planck Equation and Ito Process,” in ICML
2014
Cited alongside, same era.
Chen, C., Ding, N., and Carin, L. (2015), “On the Convergence of Stochastic Gradient MCMC Algorithms with High-order Integrators,” in NeurIPS
2015
Cited alongside, same era.
He, K., Zhang, X., Ren, S., and Sun, J. (2015), “Deep Residual Learning for Image Recognition,” CVPR
2015
Cited alongside, same era.
Ma, Y.-A., Chen, T., and Fox, E. B. (2015), “A Complete Recipe for Stochastic Gradient MCMC,” in NIPS
2015
Cited alongside, same era.
Li, C., Chen, C., Carlson, D. E., and Carin, L. (2016), “Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks,” in AAAI
2016
Cited alongside, same era.
Ruder, S. (2016), “An overview of gradient descent optimization algorithms,” CoRR
2016
Cited alongside, same era.
Raginsky, M., Rakhlin, A., and Telgarsky, M. (2017), “Non-convex Learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis,” Proceedings of Machine Learning Research
2017
Later among the works it cites.
Chen, C. (2018), “Uncertainty Estimation of Deep Neural Networks,” PhD dissertation, University of South Carolina, U.S.A
2018
Later among the works it cites.
Liang, F., Li, Q., and Zhou, L. (2018), “Bayesian Neural Networks for Selection of Drug Sensitive Genes,” Journal of the American Statistical Association
2018
Later among the works it cites.
2018
Later among the works it cites.
Xu, P., Chen, J., Zou, D., and Gu, Q. (2018), “Global Convergence of Langevin Dynamics Based Algorithms for Nonconvex Optimization,” in NeurIPS
2018
Later among the works it cites.
Deng, W., Zhang, X., Liang, F., and Lin, G. (2019), “An Adaptive Empirical Bayesian Method for Sparse Deep Learning,” in NeurIPS
2019
Later among the works it cites.
Lin, W., Khan, M. E., and Schmidt, M. (2019), “Fast and Simple Natural-Gradient Variational Inference with Mixture of Exponential-family Approximations,” in ICML
2019
Later among the works it cites.
Staib, M., Reddi, S., Kale, S., Kumar, S., and Sra, S. (2019), “Escaping saddle points with adaptive gradient methods,” in ICML
2019
Later among the works it cites.