Fetching the paper…
Reading the bibliography…
In a variety of problems originating in supervised, unsupervised, and reinforcement learning, the loss function is defined by an expectation over a collection of random variables, which might be part of a probabilistic model or the external world.
Likelihood ratio gradient estimation for stochastic systems
P. W. Glynn · 1990
Earlier work this paper cites.
Learning stochastic feedforward networks
R. M. Neal · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
R. M. Neal and G. E. Hinton · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Numerical optimization
S. J. Wright and J. Nocedal · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
Earlier work this paper cites.
Monte Carlo methods in financial engineering
P. Glasserman · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
E. Greensmith, P. L. Bartlett, and J. Baxter · 2004
Cited alongside, same era.
Gradient estimation
M. C. Fu · 2006
Cited alongside, same era.
Policy gradient in continuous time
R. Munos · 2006
Cited alongside, same era.
Evaluating derivatives: principles and techniques of algorithmic differentiation
A. Griewank and A. Walther · 2008
Cited alongside, same era.
Learning model-free robot control by a Monte Carlo EM algorithm
N. Vlassis, M. Toussaint, G. Kontes, and S. Piperidis · 2009
Cited alongside, same era.
Deep learning via Hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Black box variational inference
R. Ranganath, S. Gerrish, and D. M. Blei · 2013
Later among the works it cites.
Automated variational inference in probabilistic programming
D. Wingate and T. Weber · 2013
Later among the works it cites.
Efficient gradient-based inference through transformations between bayes nets and neural nets
D. P. Kingma and M. Welling · 2014
Later among the works it cites.
Neural variational inference and learning in belief networks
A. Mnih and K. Gregor · 2014
Later among the works it cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Wierstra, A. Förster, J. Peters, and J. Schmidhuber · 2010
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Cited alongside, same era.
K. Gregor, I. Danihelka, A. Mnih, C. Blundell, and D. Wierstra · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
Probabilistic reasoning in intelligent systems: networks of plausible inference
J. Pearl · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Later among the works it cites.
Reinforcement learning neural Turing machines
W. Zaremba and I. Sutskever · 2015
Closest in time.