Fetching the paper…
Reading the bibliography…
The posteriors over neural network weights are high dimensional and multimodal.
Diffusion for global optimization in rn
Tzuu-Shuh Chiang and Chii-Ruey Hwang · 1987
Earlier work this paper cites.
Minimum complexity density estimation
Andrew R Barron and Thomas M Cover · 1991
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Reversible jump MCMC computation and bayesian model determination
Peter J Green · 1995
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K Warmuth · 1996
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1996
Earlier work this paper cites.
Weighted csiszár-kullback-pinsker inequalities and applications to transportation inequalities
François Bolley and Cédric Villani · 2005
Earlier work this paper cites.
Estimating the integrated likelihood via posterior simulation using the harmonic mean identity
Adrian E Raftery, Michael A Newton, Jaya M Satagopan, and Pavel N Krivitsky · 2006
Earlier work this paper cites.
Construction of numerical time-average and stationary measures via Poisson equations
J. C. Mattingly, A. M. Stuart, and M. V. Tretyakov · 2010
Earlier work this paper cites.
Not MNIST Dataset
Yaroslav Bulatov · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Parallel Markov Chain Monte Carlo
Douglas N VanDerwerken and Scott C Schmidler · 2013
Earlier work this paper cites.
Distributed stochastic gradient MCMC
Sungjin Ahn, Babak Shahbaba, and Max Welling · 2014
Earlier work this paper cites.
Towards scaling up Markov Chain Monte Carlo: an adaptive subsampling approach
Rémi Bardenet, Arnaud Doucet, and Chris Holmes · 2014
Earlier work this paper cites.
Stochastic gradient Hamiltonian Monte Carlo
Tianqi Chen, Emily Fox, and Carlos Guestrin · 2014
Earlier work this paper cites.
Bayesian sampling using stochastic gradient thermostats
Nan Ding, Youhan Fang, Ryan Babbush, Changyou Chen, Robert D Skeel, and Hartmut Neven · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
The no-u-turn sampler: adaptively setting path lengths in hamiltonian monte carlo
Matthew D Hoffman and Andrew Gelman · 2014
Earlier work this paper cites.
Austerity in MCMC land: Cutting the Metropolis-Hastings budget
Anoop Korattikara, Yutian Chen, and Max Welling · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
On the convergence of stochastic gradient MCMC algorithms with high-order integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Cited alongside, same era.
A complete recipe for stochastic gradient MCMC
Yi-An Ma, Tianqi Chen, and Emily Fox · 2015
Cited alongside, same era.
Plug & play generative networks: Conditional iterative generation of images in latent space
Anh Nguyen, Jeff Clune, Yoshua Bengio, Alexey Dosovitskiy, and Jason Yosinski · 2017
Later among the works it cites.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Later among the works it cites.
Bayesian GAN
Yunus Saatchi and Andrew Gordon Wilson · 2017
Later among the works it cites.
{Euclidean, metric, and Wasserstein} gradient flows: an overview
Filippo Santambrogio · 2017
Later among the works it cites.
Super-convergence: Very fast training of residual networks using large learning rates
Leslie N Smith and Nicholay Topin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
(Non-)asymptotic properties of stochastic gradient Langevin dynamics
S. J. Vollmer, K. C. Zygalakis, and Y. W. Teh · 2015
Cited alongside, same era.
Scalable Bayesian learning of recurrent neural networks for language modeling
Zhe Gan, Chunyuan Li, Changyou Chen, Yunchen Pu, Qinliang Su, and Lawrence Carin · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2016
Cited alongside, same era.
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient Langevin dynamics
Y. Zhang, P. Liang, and M. Charikar · 2017
Later among the works it cites.
A unified particle-optimization framework for scalable bayesian sampling
Changyou Chen, Ruiyi Zhang, Wenlin Wang, Bai Li, and Liqun Chen · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of DNNs
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Later among the works it cites.
Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence Theory
Jianyi Zhang, Ruiyi Zhang, and Changyou Chen · 2018
Later among the works it cites.
User-friendly guarantees for the langevin monte carlo with inaccurate gradient
Arnak S. Dalalyan and Avetik Karagulyan · 2019
Closest in time.
Safe-bayesian generalized linear regression
Rianne de Heide, Alisa Kirichenko, Nishant Mehta, and Peter Grünwald · 2019
Closest in time.
Cyclical annealing schedule: A simple approach to mitigating KL vanishing
Hao Fu, Chunyuan Li, Xiaodong Liu, Jianfeng Gao, Asli Celikyilmaz, and Lawrence Carin · 2019
Closest in time.
Understanding mcmc dynamics as flows on the wasserstein space
Chang Liu, Jingwei Zhuo, and Jun Zhu · 2019
Closest in time.
A simple baseline for bayesian uncertainty in deep learning
Wesley J Maddox, Pavel Izmailov, Timur Garipov, Dmitry P Vetrov, and Andrew Gordon Wilson · 2019
Closest in time.
The case for bayesian deep learning
Andrew Gordon Wilson · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.