Fetching the paper…
Reading the bibliography…
To backpropagate the gradients through stochastic binary layers, we propose the augment-REINFORCE-merge (ARM) estimator that is unbiased, exhibits low variance, and has low computational complexity.
The calculation of posterior distributions by data augmentation
Martin A Tanner and Wing Hung Wong · 1987
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Peter W Glynn · 1990
Earlier work this paper cites.
Connectionist learning of belief networks
Radford M Neal · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Neural Networks for Pattern Recognition
Christopher M Bishop · 1995
Earlier work this paper cites.
The “wake-sleep” algorithm for unsupervised neural networks
Geoffrey E Hinton, Peter Dayan, Brendan J Frey, and Radford M Neal · 1995
Earlier work this paper cites.
Rao-blackwellisation of sampling schemes
George Casella and Christian P Robert · 1996
Earlier work this paper cites.
Mean field theory for sigmoid belief networks
Lawrence K Saul, Tommi Jaakkola, and Michael I Jordan · 1996
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael I Jordan, Zoubin Ghahramani, Tommi S Jaakkola, and Lawrence K Saul · 1999
Earlier work this paper cites.
The art of data augmentation
David A Van Dyk and Xiao-Li Meng · 2001
Earlier work this paper cites.
Gradient estimation
Michael C Fu · 2006
Earlier work this paper cites.
Introduction to Probability Models
Sheldon M. Ross · 2006
Earlier work this paper cites.
On the quantitative analysis of deep belief networks
Ruslan Salakhutdinov and Iain Murray · 2008
Earlier work this paper cites.
The neural autoregressive distribution estimator
Hugo Larochelle and Iain Murray · 2011
Earlier work this paper cites.
Variational Bayesian inference with stochastic search
John Paisley, David M Blei, and Michael I Jordan · 2012
Cited alongside, same era.
Negative binomial process count and mixture modeling
Mingyuan Zhou and Lawrence Carin · 2012
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Cited alongside, same era.
Karol Gregor, Ivo Danihelka, Andriy Mnih, Charles Blundell, and Daan Wierstra · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Later among the works it cites.
MuProp: Unbiased backpropagation for stochastic neural networks
Shixiang Gu, Sergey Levine, Ilya Sutskever, and Andriy Mnih · 2016
Later among the works it cites.
Variational inference for Monte Carlo objectives
Andriy Mnih and Danilo J Rezende · 2016
Later among the works it cites.
The generalized reparameterization gradient
Francisco J. R. Ruiz, Michalis K. Titsias, and David M. Blei · 2016
Later among the works it cites.
Variational inference: A review for statisticians
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Cited alongside, same era.
Monte Carlo Theory, Methods and Examples , chapter 8 Variance Reduction
Art B. Owen · 2013
Cited alongside, same era.
Learning stochastic feedforward neural networks
Yichuan Tang and Ruslan R Salakhutdinov · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Neural variational inference and learning in belief networks
Andriy Mnih and Karol Gregor · 2014
Cited alongside, same era.
Techniques for learning binary stochastic feedforward neural networks
Tapani Raiko, Mathias Berglund, Guillaume Alain, and Laurent Dinh · 2014
Cited alongside, same era.
Black box variational inference
Rajesh Ranganath, Sean Gerrish, and David Blei · 2014
Cited alongside, same era.
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Later among the works it cites.
Automatic differentiation variational inference
Alp Kucukelbir, Dustin Tran, Rajesh Ranganath, Andrew Gelman, and David M Blei · 2017
Later among the works it cites.
The Concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Later among the works it cites.
Reparameterization gradients through acceptance-rejection sampling algorithms
Christian Naesseth, Francisco Ruiz, Scott Linderman, and David Blei · 2017
Later among the works it cites.
REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J Maddison, John Lawson, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoff Roeder, and David Duvenaud · 2018
Closest in time.
Semi-implicit variational inference
Mingzhang Yin and Mingyuan Zhou · 2018
Closest in time.