Fetching the paper…
Reading the bibliography…
Recent advances in deep reinforcement learning have demonstrated the capability of learning complex control policies from many types of environments.
Risk-sensitive markov decision processes
R. A. Howard and J. E. Matheson · 1972
Earlier work this paper cites.
Markov decision processes with a new optimality criterion: Discrete time
S. C. Jaquette · 1973
Earlier work this paper cites.
The variance of discounted markov decision processes
M. Sobel · 1982
Earlier work this paper cites.
The distance between two random vectors with given dispersion matrices
I. Olkin and F. Pukelsheim · 1982
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite markov decision processes: a review
D. White · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Consideration of risk in reinforcement learning
M. Heger · 1994
Earlier work this paper cites.
Percentile performance criteria for limiting average markov decision processes
J. A. Filar, D. Krass, and K. W. Ross · 1995
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
M. L. Littman and C. Szepesvári · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Optimization of conditional value-at-risk
R. T. Rockafellar, S. Uryasev, et al · 2000
Earlier work this paper cites.
Actor-critic algorithms
V. Konda and J. Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Solving uncertain markov decision problems
J. Bagnell, A. Y. Ng, and J. Schneider · 2001
Earlier work this paper cites.
Td algorithm for the variance of return and mean-variance reinforcement learning
M. Sato, H. Kimura, and S. Kobayashi · 2001
Earlier work this paper cites.
Risk-sensitive reinforcement learning
O. Mihatsch and R. Neuneier · 2002
Earlier work this paper cites.
Q-learning for risk-sensitive control
V. S. Borkar · 2002
Earlier work this paper cites.
Reinforcement learning under circumstances beyond its control, 2003
C. Gaskett · 2003
Cited alongside, same era.
Robust control of Markov decision processes with uncertain transition matrices
A. Nilim and L. El Ghaoui · 2005
Cited alongside, same era.
Robust reinforcement learning
J. Morimoto and K. Doya · 2005
Cited alongside, same era.
Risk-sensitive reinforcement learning applied to control under constraints
P. Geibel and F. Wysotzki · 2005
Cited alongside, same era.
Risk-directed exploration in reinforcement learning
E. L. M. Law · 2005
Cited alongside, same era.
Stanley: The robot that won the darpa grand challenge
S. Thrun, M. Montemerlo, H. Dahlkamp, D. Stavens, A. Aron, J. Diebel, P. Fong, J. Gale, M. Halpenny, G. Hoffmann, et al · 2006
Cited alongside, same era.
Deep reinforcement learning in parameterized action space
M. Hausknecht and P. Stone · 2015
Later among the works it cites.
Kinematic and dynamic vehicle models for autonomous driving control design
J. Kong, M. Pfeiffer, G. Schildbach, and F. Borrelli · 2015
Later among the works it cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Y. Chow, A. Tamar, S. Mannor, and M. Pavone · 2015
Later among the works it cites.
Optimizing the cvar via sampling
A. Tamar, Y. Glassner, and S. Mannor · 2015
Later among the works it cites.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parametric return density estimation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Cited alongside, same era.
Policy gradients with variance related risk criteria
A. Tamar, D. D. Castro, and S. Mannor · 2012
Cited alongside, same era.
Smart exploration in reinforcement learning using absolute temporal difference errors
C. Gehring and D. Precup · 2013
Cited alongside, same era.
Actor-critic algorithms for risk-sensitive mdps
P. L.A. and M. Ghavamzadeh · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller · 2014
Cited alongside, same era.
Optimizing the CVaR via Sampling
A. Tamar, Y. Glassner, and S. Mannor · 2014
Cited alongside, same era.
A. Tamar, D. D. Castro, and S. Mannor · 2016
Later among the works it cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Later among the works it cites.
On a formal model of safe and scalable self-driving cars
S. Shalev-Shwartz, S. Shammah, and A. Shashua · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
CARLA: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Later among the works it cites.
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2017
Later among the works it cites.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, A. Muldal, N. Heess, and T. Lillicrap · 2018
Later among the works it cites.
Implicit quantile networks for distributional reinforcement learning
W. Dabney, G. Ostrovski, D. Silver, and R. Munos · 2018
Later among the works it cites.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. M. Czarnecki, and et al · 2019
Closest in time.