Fetching the paper…
Reading the bibliography…
Stochastic gradient Markov chain Monte Carlo (SG-MCMC) has become increasingly popular for simulating posterior samples in large-scale Bayesian modeling.
Hybrid Monte Carlo
Simon Duane, Anthony D Kennedy, Brian J Pendleton, and Duncan Roweth · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Meta-neural networks that learn by learning
Devang K Naik and RJ Mammone · 1992
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael I Jordan, Zoubin Ghahramani, Tommi S Jaakkola, and Lawrence K Saul · 1999
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Matthew James Beal · 2003
Earlier work this paper cites.
Slice sampling
Radford M Neal · 2003
Earlier work this paper cites.
Existence and construction of dynamical potential in nonequilibrium processes without detailed balance
Lan Yin and Ping Ao · 2006
Earlier work this paper cites.
Eric Brochu, Vlad M Cora, and Nando De Freitas · 2010
Earlier work this paper cites.
Adaptively scaling the Metropolis algorithm using expected squared jumped distance
Cristian Pasarica and Andrew Gelman · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Riemann manifold Langevin and Hamiltonian Monte Carlo methods
Mark Girolami and Ben Calderhead · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian posterior sampling via stochastic gradient fisher scoring
Sungjin Ahn, Anoop Korattikara, and Max Welling · 2012
Earlier work this paper cites.
Relation of a new interpretation of stochastic differential equations to ito process
Jianghong Shi, Tianqi Chen, Ruoshi Yuan, Bo Yuan, and Ping Ao · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Stochastic gradient Riemannian Langevin dynamics on the probability simplex
Sam Patterson and Yee Whye Teh · 2013
Cited alongside, same era.
Stochastic gradient Hamiltonian Monte Carlo
Tianqi Chen, Emily Fox, and Carlos Guestrin · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Bayesian sampling using stochastic gradient thermostats
Nan Ding, Youhan Fang, Ryan Babbush, Changyou Chen, Robert D Skeel, and Hartmut Neven · 2014
Cited alongside, same era.
NICE: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Cited alongside, same era.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Later among the works it cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Later among the works it cites.
Learning to learn without gradient descent by gradient descent
Yutian Chen, Matthew W Hoffman, Sergio Gómez Colmenarejo, Misha Denil, Timothy P Lillicrap, Matt Botvinick, and Nando Freitas · 2017
Later among the works it cites.
Learning and policy search in stochastic dynamical systems with Bayesian neural networks
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft · 2017
Later among the works it cites.
Density estimation using Real NVP
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Black box variational inference
Rajesh Ranganath, Sean Gerrish, and David Blei · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Cited alongside, same era.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Cited alongside, same era.
A complete recipe for stochastic gradient MCMC
Yi-An Ma, Tianqi Chen, and Emily Fox · 2015
Cited alongside, same era.
Markov chain Monte Carlo and variational inference: Bridging the gap
Tim Salimans, Diederik Kingma, and Max Welling · 2015
Cited alongside, same era.
Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner · 2017
Later among the works it cites.
Model-Agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Scalable Bayesian learning of Recurrent neural networks for language modeling
Zhe Gan, Chunyuan Li, Changyou Chen, Yunchen Pu, Qinliang Su, and Lawrence Carin · 2017
Later among the works it cites.
Learning to optimize
Ke Li and Jitendra Malik · 2017
Later among the works it cites.
Dropout inference in Bayesian neural networks with Alpha-divergences
Yingzhen Li and Yarin Gal · 2017
Later among the works it cites.
Multiplicative normalizing flows for variational Bayesian neural networks
Christos Louizos and Max Welling · 2017
Later among the works it cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2017
Later among the works it cites.
A-NICE-MC: Adversarial training for MCMC
Jiaming Song, Shengjia Zhao, and Stefano Ermon · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gómez Colmenarejo, Misha Denil, Nando Freitas, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Generalizing Hamiltonian Monte Carlo with neural networks
Daniel Levy, Matt D. Hoffman, and Jascha Sohl-Dickstein · 2018
Closest in time.
Gradient estimators for implicit models
Yingzhen Li and Richard E. Turner · 2018
Closest in time.
Variational continual learning
Cuong V. Nguyen, Yingzhen Li, Thang D. Bui, and Richard E. Turner · 2018
Closest in time.
A scalable Laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Closest in time.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Closest in time.