Fetching the paper…
Reading the bibliography…
In this paper we propose a general framework of performing MCMC with only a mini-batch of data.
“A stochastic approximation method”
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
“Probability inequalities for sums of bounded random variables”
Wassily Hoeffding · 1963
Earlier work this paper cites.
“Monte carlo calculations of the radial distribution functions for a proton? electron plasma”
Av Barker · 1965
Earlier work this paper cites.
“A bound for the error in the normal approximation to the distribution of a sum of dependent random variables”
Charles Stein · 1972
Earlier work this paper cites.
“Probability inequalities for the sum in sampling without replacement”
Robert Serfling · 1974
Earlier work this paper cites.
“Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images”
Stuart Geman and Donald Geman · 1984
Earlier work this paper cites.
“Hybrid monte carlo”
Simon Duane, Anthony Kennedy, Brian Pendleton and Duncan Roweth · 1987
Earlier work this paper cites.
“Markov chain Monte Carlo maximum likelihood”
Charles Geyer · 1991
Earlier work this paper cites.
“Multicanonical ensemble: A new approach to simulate first-order phase transitions”
Bernd Berg and Thomas Neuhaus · 1992
Earlier work this paper cites.
“Exponential convergence of Langevin distributions and their discrete approximations”
Gareth Roberts and Richard Tweedie · 1996
Earlier work this paper cites.
“Optimal scaling for various Metropolis-Hastings algorithms”
Gareth Roberts and Jeffrey Rosenthal · 2001
Earlier work this paper cites.
“Latent dirichlet allocation”
David Blei, Andrew Ng and Michael Jordan · 2003
Earlier work this paper cites.
“Ensemble selection from libraries of models”
Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew and Alex Ksikes · 2004
Earlier work this paper cites.
“The author-topic model for authors and documents”
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers and Padhraic Smyth · 2004
Earlier work this paper cites.
“Parallel tempering: Theory, applications, and new perspectives”
D.. Earl and M.. Deem · 2005
Earlier work this paper cites.
“Sampling for Bayesian computation with large datasets”, 2005
Zaijing Huang and Andrew Gelman · 2005
Earlier work this paper cites.
“On self-normalized sums and Student’s statistic”
Sergei’evich Novak · 2005
Earlier work this paper cites.
“Discussion paper equi-energy sampler with applications in statistical inference and statistical mechanics”
SC Kou, Qing Zhou and Wing Wong · 2006
Earlier work this paper cites.
“Visualizing data using t-SNE”
Laurens van Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
“Reconstructing the energy landscape of a distribution from Monte Carlo samples”
Qing Zhou and Wing Wong · 2008
Cited alongside, same era.
“The pseudo-marginal approach for efficient Monte Carlo computations”
Christophe Andrieu and Gareth Roberts · 2009
Cited alongside, same era.
“Robust optimization”
Aharon Ben-Tal, Laurent El and Arkadi Nemirovski · 2009
Cited alongside, same era.
“Robust stochastic approximation approach to stochastic programming”
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan and Alexander Shapiro · 2009
Cited alongside, same era.
“Understanding the difficulty of training deep feedforward neural networks.”
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
“MNIST handwritten digit database”
Yann LeCun, Corinna Cortes and Christopher Burges · 2010
Cited alongside, same era.
“Firefly Monte Carlo: Exact MCMC with subsets of data”
Dougal Maclaurin and Ryan Adams · 2014
Later among the works it cites.
“Scalable and robust Bayesian inference via the median posterior”
Stanislav Minsker, Sanvesh Srivastava, Lizhen Lin and David Dunson · 2014
Later among the works it cites.
“On Markov chain Monte Carlo methods for tall data”
R\’emi Bardenet, Arnaud Doucet and Chris Holmes · 2015
Later among the works it cites.
“The fundamental incompatibility of scalable Hamiltonian Monte Carlo and naive data subsampling”
Michael Betancourt · 2015
Later among the works it cites.
“The Loss Surfaces of Multilayer Networks.”
Anna Choromanska et al · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Adaptive subgradient methods for online learning and stochastic optimization”
John Duchi, Elad Hazan and Yoram Singer · 2011
Cited alongside, same era.
“MCMC using Hamiltonian dynamics”
Radford Neal · 2011
Cited alongside, same era.
“Bayesian learning via stochastic gradient Langevin dynamics”
Max Welling and Yee Teh · 2011
Cited alongside, same era.
“Bayesian Posterior Sampling via Stochastic Gradient Fisher Scoring.”
Sungjin Ahn, Anoop Balan and Max Welling · 2012
Cited alongside, same era.
“Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude”
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
“Asymptotically exact, embarrassingly parallel MCMC”
Willie Neiswanger, Chong Wang and Eric Xing · 2013
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Later among the works it cites.
“Approximations of markov chains and bayesian inference”
James Johndrow, Jonathan Mattingly, Sayan Mukherjee and David Dunson · 2015
Later among the works it cites.
Dmytro Mishkin and Jiri Matas · 2015
Later among the works it cites.
“WASP: Scalable Bayes via barycenters of subset posteriors”
Sanvesh Srivastava, Volkan Cevher, Quoc Dinh and David Dunson · 2015
Later among the works it cites.
“Bayesian fractional posteriors”
Anirban Bhattacharya, Debdeep Pati and Yun Yang · 2016
Later among the works it cites.
“Bridging the gap between stochastic gradient MCMC and stochastic optimization”
Changyou Chen et al · 2016
Later among the works it cites.
“Deep learning without poor local minima”
Kenji Kawaguchi · 2016
Later among the works it cites.
“On large-batch training for deep learning: Generalization gap and sharp minima”
Nitish Keskar et al · 2016
Later among the works it cites.
“Bayes and big data: The consensus Monte Carlo algorithm”
Steven Scott et al · 2016
Later among the works it cites.
“An Efficient Minibatch Acceptance Test for Metropolis-Hastings”
Daniel Seita, Xinlei Pan, Haoyu Chen and John Canny · 2016
Later among the works it cites.
“Consistency and fluctuations for stochastic gradient Langevin dynamics”
Yee Teh, Alexandre Thiery and Sebastian Vollmer · 2016
Later among the works it cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang et al · 2016
Later among the works it cites.
“Snapshot Ensembles: Train 1, get M for free”
Gao Huang et al · 2017
Closest in time.