Fetching the paper…
Reading the bibliography…
Recently, there has been a surge in interest in safe and robust techniques within reinforcement learning (RL).
Stochastic games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Game variant of a problem on optimal stopping. In Soviet Math. Dokl
EB Dynkin. 1967 · 1967
Earlier work this paper cites.
The big match
David Blackwell and Tom S Ferguson. 1968 · 1968
Earlier work this paper cites.
On stochastic games
A Maitra and T Parthasarathy. 1970 · 1970
Earlier work this paper cites.
On the existence of good Markov strategies
Theodore Preston Hill. 1979 · 1979
Earlier work this paper cites.
Nonlinear programming and stationary equilibria in stochastic games
Jerzy A Filar, Todd A Schultz, Frank Thuijsman, and OJ Vrieze. 1991 · 1991
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Optimal stopping, free boundary, and American option in a jump-diffusion model
Huyên Pham. 1997 · 1997
Earlier work this paper cites.
Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy. 1999 · 1999
Earlier work this paper cites.
Robust reinforcement learning. In Advances in Neural Information Processing Systems
Jun Morimoto and Kenji Doya. 2001 · 2001
Earlier work this paper cites.
Combined stochastic control and optimal stopping, and application to numerical approximation of combined stochastic and impulse control
Jean-Philippe Chancelier, Bernt Øksendal, and Agnès Sulem. 2002 · 2002
Earlier work this paper cites.
Optimal portfolios under a value-at-risk constraint
Ka-Fai Cedric Yiu. 2004 · 2004
Earlier work this paper cites.
Optimal investment strategy to minimize the probability of lifetime ruin
Virginia R Young. 2004 · 2004
Cited alongside, same era.
A note on sufficient conditions for no arbitrage
Peter Carr and Dilip B Madan. 2005 · 2005
Cited alongside, same era.
A homogeneous mobile robot team that is fault-tolerant
Toshiyuki Yasuda, Kazuhiro Ohkura, and Kanji Ueda. 2006 · 2006
Cited alongside, same era.
Stochastic games of control and stopping for a linear diffusion
Ioannis Karatzas and William Sudderth. 2006 · 2006
Cited alongside, same era.
Optimal stopping and free-boundary problems
Goran Peskir and Albert Shiryaev. 2006 · 2006
Cited alongside, same era.
Algorithmic game theory
Noam Nisan, Tim Roughgarden, Eva Tardos, and Vijay V Vazirani. 2007 · 2007
Cited alongside, same era.
Markets contagion during financial crisis: A regime-switching approach
Feng Guo, Carl R Chen, and Ying Sophie Huang. 2011 · 2011
Later among the works it cites.
Optimal investment with counterparty risk: a default-density model approach
Ying Jiao and Huyên Pham. 2011 · 2011
Later among the works it cites.
Safe exploration of state and action spaces in reinforcement learning
Javier Garcia and Fernando Fernández. 2012 · 2012
Later among the works it cites.
Optimal stopping and stochastic control differential games for jump diffusions
Fouzia Baghery, Sven Haadem, Bernt Øksendal, and Isabelle Turpin. 2013 · 2013
Later among the works it cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Later among the works it cites.
Interim monitoring of clinical trials: Decision theory, dynamic programming and optimal stopping
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Approximate dynamic programming
Dimitri P Bertsekas. 2008 · 2008
Cited alongside, same era.
Reinforcement learning–based fault-tolerant control with application to flux cored wire system
Dapeng Zhang and Zhiwei Gao. 2018 · 2009
Cited alongside, same era.
Enhancing R&D in science-based industry: An optimal stopping model for drug discovery
Guozhen Zhao and Wen Chen. 2009 · 2009
Cited alongside, same era.
Autonomous helicopter aerobatics through apprenticeship learning
Pieter Abbeel, Adam Coates, and Andrew Y Ng. 2010 · 2010
Cited alongside, same era.
Reinforcement learning-based multi-agent system for network traffic signal control
Itamar Arel, Cong Liu, T Urbanik, and AG Kohls. 2010 · 2010
Cited alongside, same era.
Minimizing the probability of lifetime ruin under stochastic volatility
Erhan Bayraktar, Xueying Hu, and Virginia R Young. 2011 · 2011
Cited alongside, same era.
Christopher Jennison and Bruce W Turnbull. 2013 · 2013
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández. 2015 · 2015
Later among the works it cites.
Optimal stopping with private information
Thomas Kruse and Philipp Strack. 2015 · 2015
Later among the works it cites.
Policy gradient for coherent risk measures. In Advances in Neural Information Processing Systems
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor. 2015 · 2015
Later among the works it cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua. 2016 · 2016
Later among the works it cites.
David Mguni. 2018 · 2018
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi. 2019 · 2019
Closest in time.