Fetching the paper…
Reading the bibliography…
Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult.
Parametric inference for imperfectly observed Gibbsian fields
Younes, Laurent · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Exchange Monte Carlo method and application to spin glass simulations
Hukushima, Koji and Nemoto, Koji · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Extended ensemble Monte Carlo
Iba, Yukito · 2001
Earlier work this paper cites.
On the quantitative analysis of deep belief networks
Salakhutdinov, Ruslan and Murray, Iain · 2008
Earlier work this paper cites.
Training restricted Boltzmann machines using approximations to the likelihood gradient
Tieleman, Tijmen · 2008
Earlier work this paper cites.
Inductive principles for restricted Boltzmann machine learning
Marlin, Benjamin, Swersky, Kevin, Chen, Bo, and Freitas, Nando · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Yoshua, Léonard, Nicholas, and Courville, Aaron · 2013
Earlier work this paper cites.
Gregor, Karol, Danihelka, Ivo, Mnih, Andriy, Blundell, Charles, and Wierstra, Daan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, Diederik P and Welling, Max · 2014
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Kingma, Diederik P, Mohamed, Shakir, Rezende, Danilo Jimenez, and Welling, Max · 2014
Earlier work this paper cites.
Neural variational inference and learning in belief networks
Mnih, Andriy and Gregor, Karol · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, Danilo Jimenez, Mohamed, Shakir, and Wierstra, Daan · 2014
Cited alongside, same era.
Importance weighted autoencoders
Burda, Yuri, Grosse, Roger, and Salakhutdinov, Ruslan · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, Djork-Arné, Unterthiner, Thomas, and Hochreiter, Sepp · 2015
Cited alongside, same era.
Deep generative image models using a Laplacian pyramid of adversarial networks
Denton, Emily L, Chintala, Soumith, Fergus, Rob, et al · 2015
Cited alongside, same era.
Muprop: Unbiased backpropagation for stochastic neural networks
Gu, Shixiang, Levine, Sergey, Sutskever, Ilya, and Mnih, Andriy · 2015
Improved variational inference with inverse autoregressive flow
Kingma, Diederik P, Salimans, Tim, Jozefowicz, Rafal, Chen, Xi, Sutskever, Ilya, and Welling, Max · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, Chris J, Mnih, Andriy, and Teh, Yee Whye · 2016
Later among the works it cites.
Variational inference for Monte Carlo objectives
Mnih, Andriy and Rezende, Danilo · 2016
Later among the works it cites.
Discrete variational autoencoders
Rolfe, Jason Tyler · 2016
Later among the works it cites.
Ladder variational autoencoders
Sønderby, Casper Kaae, Raiko, Tapani, Maaløe, Lars, Sønderby, Søren Kaae, and Winther, Ole · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Lake, Brenden M, Salakhutdinov, Ruslan, and Tenenbaum, Joshua B · 2015
Cited alongside, same era.
Makhzani, Alireza, Shlens, Jonathon, Jaitly, Navdeep, Goodfellow, Ian, and Frey, Brendan · 2015
Cited alongside, same era.
Generating sentences from a continuous space
Bowman, Samuel R, Vilnis, Luke, Vinyals, Oriol, Dai, Andrew, Jozefowicz, Rafal, and Bengio, Samy · 2016
Cited alongside, same era.
Chen, Xi, Kingma, Diederik P, Salimans, Tim, Duan, Yan, Dhariwal, Prafulla, Schulman, John, Sutskever, Ilya, and Abbeel, Pieter · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Chung, Junyoung, Ahn, Sungjin, and Bengio, Yoshua · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Cited alongside, same era.
Van Den Oord, Aäron, Kalchbrenner, Nal, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Squeeze-and-excitation networks
Hu, Jie, Shen, Li, and Sun, Gang · 2017
Later among the works it cites.
Self-normalizing neural networks
Klambauer, Günter, Unterthiner, Thomas, Mayr, Andreas, and Hochreiter, Sepp · 2017
Later among the works it cites.
Semi-supervised generation with cluster-aware generative models
Maaløe, Lars, Fraccaro, Marco, and Winther, Ole · 2017
Later among the works it cites.
Parallel multiscale autoregressive density estimation
Reed, Scott E, van den Oord, Aäron, Kalchbrenner, Nal, Gómez, Sergio, Wang, Ziyu, Belov, Dan, and de Freitas, Nando · 2017
Later among the works it cites.
Salimans, Tim, Karpathy, Andrej, Chen, Xi, and Kingma, Diederik P · 2017
Later among the works it cites.
Tomczak, Jakub M and Welling, Max · 2017
Later among the works it cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, George, Mnih, Andriy, Maddison, Chris J, Lawson, John, and Sohl-Dickstein, Jascha · 2017
Later among the works it cites.