Fetching the paper…
Reading the bibliography…
In reinforcement learning, an agent attempts to learn high-performing behaviors through interacting with the environment, such behaviors are often quantified in the form of a reward function.
Linear Programming and Finite Markovian Control Problems
L. Kallenberg · 1983
Earlier work this paper cites.
Optimal policies for controlled markov chains with a constraint
F. J. Beutler and K. W. Ross · 1985
Earlier work this paper cites.
Constrained markov decision processes with queueing applications
K. W. Ross · 1985
Earlier work this paper cites.
Time-average optimal constrained semi-markov decision processes
F. J. Beutler and K. W. Ross · 1986
Earlier work this paper cites.
Randomized and past-dependent policies for markov decision processes with multiple constraints
K. W. Ross · 1989
Earlier work this paper cites.
Markov decision processes with sample path constraints: the communicating case
K. W. Ross and R. Varadarajan · 1989
Earlier work this paper cites.
Multichain markov decision processes with a sample path constraint: A decomposition approach
K. W. Ross and R. Varadarajan · 1991
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Computational science and engineering
G. Strang · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mulling, and Y. Altun · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Safe policy iteration
M. Pirotta, M. Restelli, A. Pecorino, and D. Calandriello · 2013
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
D. P. Bertsekas · 2014
Cited alongside, same era.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Guided policy search via approximate mirror descent
W. H. Montgomery and S. Levine · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control
Y. Chow, O. Nachum, A. Faust, M. Ghavamzadeh, and E. Duenez-Guzman · 2019
Later among the works it cites.
A survey of reinforcement learning informed by natural language
J. Luketina, N. Nardelli, G. Farquhar, J. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
A. Ray, J. Achiam, and D. Amodei · 2019
Later among the works it cites.
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2019
Later among the works it cites.
Supervised policy update for deep reinforcement learning
Q. Vuong, Y. Zhang, and K. W. Ross · 2019
Later among the works it cites.
NeurIPS 2018 Invited Talk: Reproducible, Reusable, and Robust Reinforcement Learning, 2018
J. Pineau · 2020
Closest in time.
Responsive safety in reinforcement learning by pid lagrangian methods
A. Stooke, J. Achiam, and P. Abbeel · 2020
Closest in time.
Projection-based constrained policy optimization
T.-Y. Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge · 2020
Closest in time.