Fetching the paper…
Reading the bibliography…
The policy gradients of the expected return objective can react slowly to rare rewards.
Theory of games and economic behavior
John Von Neumann and Oskar Morgenstern · 1953
Earlier work this paper cites.
Risk aversion in the small and in the large
John W Pratt · 1964
Earlier work this paper cites.
Risk-sensitive Markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Essays in the theory of risk-bearing
Kenneth Joseph Arrow · 1974
Earlier work this paper cites.
Logarithmic transformations for discrete-time, finite-horizon stochastic control problems
Francesca Albertini and Wolfgang J Runggaldier · 1988
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Risk-sensitive planning with probabilistic decision graphs
Sven Koenig and Reid G Simmons · 1994
Earlier work this paper cites.
Optimal control of Markov decision processes for performance and robustness
Stefano Coraluppi · 1997
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1997
Earlier work this paper cites.
Risk sensitive markov decision processes
Steven I Marcus, Emmanual Fernández-Gaucherand, Daniel Hernández-Hernandez, Stefano Coraluppi, and Pedram Fard · 1997
Earlier work this paper cites.
Risk sensitive reinforcement learning
Ralph Neuneier and Oliver Mihatsch · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Cited alongside, same era.
Risk-sensitive optimal control for markov decision processes with monotone cost
Vivek S Borkar and Sean P Meyn · 2002
Cited alongside, same era.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Cited alongside, same era.
Feynman-kac formulae
Pierre Del Moral · 2004
Cited alongside, same era.
Linear theory for control of nonlinear stochastic systems
Hilbert J Kappen · 2005
Cited alongside, same era.
Linearly-solvable markov decision problems
Emanuel Todorov · 2006
Cited alongside, same era.
Information theory of decisions and actions
Naftali Tishby and Daniel Polani · 2011
Later among the works it cites.
Risk sensitive path integral control
Bart van den Broek, Wim Wiegerinck, and Hilbert Kappen · 2012
Later among the works it cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Later among the works it cites.
On some properties of markov chain monte carlo simulation methods based on the particle filter
Michael K Pitt, Ralph dos Santos Silva, Paolo Giordani, and Robert Kohn · 2012
Later among the works it cites.
More risk-sensitive markov decision processes
Nicole Bäuerle and Ulrich Rieder · 2013
Later among the works it cites.
Risk-sensitive reinforcement learning
Yun Shen, Michael J Tobia, Tobias Sommer, and Klaus Obermayer · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probabilistic inference for solving discrete and continuous state markov decision processes
Marc Toussaint and Amos Storkey · 2006
Cited alongside, same era.
On solving general state-space sequential decision problems using inference algorithms
Matt Hoffman, Arnaud Doucet, Nando De Freitas, and Ajay Jasra · 2007
Cited alongside, same era.
Sequential decision making in general state space models
Nikolas Kantas · 2009
Cited alongside, same era.
An approximate inference approach to temporal optimization in optimal control
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2010
Cited alongside, same era.
A tutorial on particle filtering and smoothing: fiteen years later
Arnaud Doucet and Adam M Johansen · 2011
Cited alongside, same era.
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Later among the works it cites.
On particle methods for parameter estimation in state-space models
Nikolas Kantas, Arnaud Doucet, Sumeetpal S Singh, Jan Maciejowski, Nicolas Chopin, et al · 2015
Later among the works it cites.
Importance Weighted Autoencoder
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov · 2016
Later among the works it cites.
Variational Inference for Monte Carlo Objectives
Andriy Mnih and Danilo Rezende · 2016
Later among the works it cites.
Particle smoothing for hidden diffusion processes: Adaptive path integral smoother
H-Ch Ruiz and HJ Kappen · 2016
Later among the works it cites.