Fetching the paper…
Reading the bibliography…
Deploying Reinforcement Learning (RL) agents to solve real-world applications often requires satisfying complex system constraints.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
An empirical investigation of the challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization, 2017
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Earlier work this paper cites.
A review on bilevel optimization: from classical to evolutionary approaches and applications
Ankur Sinha, Pekka Malo, and Kalyanmoy Deb · 2017
Earlier work this paper cites.
A deep hierarchical approach to lifelong learning in minecraft
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J Mankowitz, and Shie Mannor · 2017
Cited alongside, same era.
On ensuring that intelligent machines are well-behaved
Philip S Thomas, Bruno Castro da Silva, Andrew G Barto, and Emma Brunskill · 2017
Cited alongside, same era.
Relative entropy regularized policy iteration
Abbas Abdolmaleki, Jost Tobias Springenberg, Jonas Degrave, Steven Bohez, Yuval Tassa, Dan Belov, Nicolas Heess, and Martin A. Riedmiller · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning, 2018
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Value constrained model-free continuous control
Steven Bohez, Abbas Abdolmaleki, Michael Neunert, Jonas Buchli, Nicolas Heess, and Raia Hadsell · 2019
Later among the works it cites.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel J. Mankowitz, and Todd Hester · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Meta-Gradient Reinforcement Learning
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
Metatrace actor-critic: Online step-size tuning by meta-gradient descent for reinforcement learning control, 2018
Kenny Young, Baoxiang Wang, and Matthew E. Taylor · 2018
Cited alongside, same era.
An empirical investigation of the challenges of real-world reinforcement learning, 2020b
Gabriel Dulac-Arnold, Nir Levine, Daniel J. Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester
Cited in the paper.
Discovery of useful questions as auxiliary tasks
Vivek Veeriah, Matteo Hessel, Zhongwen Xu, Richard Lewis, Janarthanan Rajendran, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2019
Later among the works it cites.
Exploration-exploitation in constrained mdps, 2020
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Closest in time.
Constrained markov decision processes via backward value functions
Harsh Satija, Philip Amortila, and Joelle Pineau · 2020
Closest in time.
Self-Tuning Deep Reinforcement Learning
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Closest in time.
Reward constrained interactive recommendation with natural language feedback, 2020
Ruiyi Zhang, Tong Yu, Yilin Shen, Hongxia Jin, Changyou Chen, and Lawrence Carin · 2020
Closest in time.