Fetching the paper…
Reading the bibliography…
Many potential applications of reinforcement learning (RL) require guarantees that the agent will perform well in the face of disturbances to the dynamics or reward function.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Convex Analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Robust and optimal control , volume 40
Kemin Zhou, John Comstock Doyle, Keith Glover, et al · 1996
Earlier work this paper cites.
Solving uncertain Markov decision processes
J Andrew Bagnell, Andrew Y Ng, and Jeff G Schneider · 2001
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Earlier work this paper cites.
Robustness in markov decision problems with uncertain transition matrices
Arnab Nilim and Laurent Ghaoui · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Game theory, maximum entropy, minimum discrepancy and robust bayesian decision theory
Peter D Grünwald, A Philip Dawid, et al · 2004
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Hilbert J Kappen · 2005
Earlier work this paper cites.
Robust reinforcement learning
Jun Morimoto and Kenji Doya · 2005
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
Maximum causal entropy correlated equilibria for markov games
Brian D Ziebart, J Andrew Bagnell, and Anind K Dey · 2011
Earlier work this paper cites.
Transfer in reinforcement learning: a framework and a survey
Alessandro Lazaric · 2012
Earlier work this paper cites.
Feedback control theory
John C Doyle, Bruce A Francis, and Allen R Tannenbaum · 2013
Earlier work this paper cites.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Faulty reward functions in the wild
Jack Clark and Dario Amodei · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
The CMA evolution strategy: A tutorial
Nikolaus Hansen · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2016
Cited alongside, same era.
CAD2RL: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2016
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2019
Later among the works it cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2019
Later among the works it cites.
Learning to walk via deep reinforcement learning
Tuomas Haarnoja, Sehoon Ha, Aurick Zhou, Jie Tan, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Svqn: Sequential variational soft q-learning networks
Shiyu Huang, Hang Su, Jun Zhu, and Ting Chen · 2019
Later among the works it cites.
Adversarial examples are not bugs, they are features
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Later among the works it cites.
Tsallis reinforcement learning: A unified framework for maximum entropy reinforcement learning
Kyungjae Lee, Sungyub Kim, Sungbin Lim, Sungjoon Choi, and Songhwai Oh · 2019
Later among the works it cites.
Beyond confidence regions: Tight bayesian ambiguity sets for robust MDPs
Reazul Hasan Russel and Marek Petrik · 2019
Later among the works it cites.
Action robust reinforcement learning and applications in continuous control
Chen Tessler, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
Positive-unlabeled reward learning
Danfei Xu and Misha Denil · 2019
Later among the works it cites.
Less is more: Rethinking probabilistic models of human behavior
Andreea Bobu, Dexter RR Scobee, Jaime F Fisac, S Shankar Sastry, and Anca D Dragan · 2020
Later among the works it cites.
Robust reinforcement learning via adversarial training with langevin dynamics
Parameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland, Cheng Shi, and Volkan Cevher · 2020
Later among the works it cites.
Lipschitz lifelong reinforcement learning
Erwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai, Emmanuel Rachelson, and Michael L Littman · 2020
Later among the works it cites.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Later among the works it cites.
Entropic risk constrained soft-robust policy optimization
Reazul Hasan Russel, Bahram Behzadian, and Marek Petrik · 2020
Later among the works it cites.
Worst cases policy gradients
Yichuan Charlie Tang, Jian Zhang, and Ruslan Salakhutdinov · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Twice regularized MDPs and the equivalence between robustness and regularization
Esther Derman, Matthieu Geist, and Shie Mannor · 2021
Closest in time.
Reazul Hasan Russel, Mouhacine Benosman, Jeroen Van Baar, and Radu Corcodel · 2021
Closest in time.
Recovery rl: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg · 2021
Closest in time.