Fetching the paper…
Reading the bibliography…
A trustworthy reinforcement learning algorithm should be competent in solving challenging real-world problems, including {robustly} handling uncertainties, satisfying {safety} constraints to avoid catastrophic failures, and {generalizing} to unseen scenarios during deployments.
Evaluation uncertainty in data-driven self-driving testing. In 2019 IEEE Intelligent Transportation Systems Conference (ITSC) . IEEE, 1902–1907
Zhiyuan Huang, Mansur Arief, Henry Lam, and Ding Zhao. 2019 · 1907
Earlier work this paper cites.
Theory of games and economic behavior
Oskar Morgenstern and John Von Neumann. 1953 · 1953
Earlier work this paper cites.
Learning to Achieve Goals. In IN PROC. OF IJCAI-93 . Morgan Kaufmann, 1094–1098
Leslie Pack Kaelbling. 1993 · 1993
Earlier work this paper cites.
Experimental evidence on players’ models of other players
Dale O Stahl II and Paul W Wilson. 1994 · 1994
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M Mitchell. 1995 · 1995
Earlier work this paper cites.
Evolutionary robotics and the radical envelope-of-noise hypothesis
Nick Jakobi. 1997 · 1997
Earlier work this paper cites.
Humans and automation: Use, misuse, disuse, abuse
Raja Parasuraman and Victor Riley. 1997 · 1997
Earlier work this paper cites.
Constrained Markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Eitan Altman. 1998 · 1998
Earlier work this paper cites.
Reinforcement learning - an introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Bayesian domain randomization for sim-to-real transfer
Fabio Muratore, Christian Eilers, Michael Gienger, and Jan Peters. 2020 · 2003
Earlier work this paper cites.
Trust in automation: Designing for appropriate reliance
John D Lee and Katrina A See. 2004 · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Vivek S Borkar. 2005 · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui. 2005 · 2005
Earlier work this paper cites.
The robustness-performance tradeoff in Markov decision processes
Huan Xu and Shie Mannor. 2006 · 2006
Earlier work this paper cites.
Goal inference as inverse planning. In Proceedings of the Annual Meeting of the Cognitive Science Society , Vol. 29
Chris L Baker, Joshua B Tenenbaum, and Rebecca R Saxe. 2007 · 2007
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach. In Proceedings of the 24th international conference on Machine learning . 1015–1022
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli. 2007 · 2007
Earlier work this paper cites.
Privacy-preserving reinforcement learning. In Proceedings of the 25th international conference on Machine learning . 864–871
Jun Sakuma, Shigenobu Kobayashi, and Rebecca N Wright. 2008 · 2008
Earlier work this paper cites.
Value-Based Policy Teaching with Active Indirect Elicitation.. In AAAI , Vol. 8. 208–214
Haoqi Zhang and David C Parkes. 2008 · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning.. In Aaai , Vol. 8. Chicago, IL, USA, 1433–1438
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
Paweł Wawrzyński. 2009 · 2009
Earlier work this paper cites.
Distributionally Robust Markov Decision Processes.. In NIPS . 2505–2513
Huan Xu and Shie Mannor. 2010 · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan. 2012 · 2012
Earlier work this paper cites.
Infinite-horizon model predictive control for periodic tasks with contacts
Tom Erez, Yuval Tassa, and Emanuel Todorov. 2012 · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor. 2012 · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Earlier work this paper cites.
Learning an internal dynamics model from control demonstration. In International Conference on Machine Learning . PMLR, 606–614
Matthew Golub, Steven Chase, and Byron Yu. 2013 · 2013
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Robust Markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem. 2013 · 2013
Earlier work this paper cites.
Online multi-task learning for policy gradient methods. In International conference on machine learning . PMLR, 1206–1214
Haitham Bou Ammar, Eric Eaton, Paul Ruvolo, and Matthew Taylor. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Towards deep neural network architectures robust to adversarial examples
Shixiang Gu and Luca Rigazio. 2014 · 2014
Earlier work this paper cites.
Social learning theory
Ronald L Akers and Wesley G Jennings. 2015 · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández. 2015 · 2015
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor. 2015 · 2015
Earlier work this paper cites.
Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids. In 2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 5307–5314
Igor Mordatch, Kendall Lowrey, and Emanuel Todorov. 2015 · 2015
Earlier work this paper cites.
Robust partially observable Markov decision process. In International Conference on Machine Learning . PMLR, 106–115
Takayuki Osogami. 2015 · 2015
Earlier work this paper cites.
Universal Value Function Approximators. In Proceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37) , Francis Bach and David Blei (Eds.). PMLR, Lille, France, 1312–1320
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. 2015 · 2015
Earlier work this paper cites.
Trust region policy optimization. In International conference on machine learning . PMLR, 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Earlier work this paper cites.
Optimizing the CVaR via sampling. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 29
Aviv Tamar, Yonatan Glassner, and Shie Mannor. 2015 · 2015
Earlier work this paper cites.
Distributionally robust counterpart in Markov decision processes
Pengqian Yu and Huan Xu. 2015 · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Robust MDPs with k-rectangular uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu. 2016 · 2016
Earlier work this paper cites.
Epopt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine. 2016 · 2016
Earlier work this paper cites.
Cad2rl: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine. 2016 · 2016
Earlier work this paper cites.
Constrained policy optimization. In International Conference on Machine Learning . PMLR, 22–31
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. 2017 · 2017
Earlier work this paper cites.
Reinforcement learning for pivoting task
Rika Antonova, Silvia Cruciani, Christian Smith, and Danica Kragic. 2017 · 2017
Earlier work this paper cites.
Whatever does not kill deep reinforcement learning, makes it stronger
Vahid Behzadan and Arslan Munir. 2017b · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P Schoellig, and Andreas Krause. 2017 · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp) . IEEE, 39–57
Nicholas Carlini and David Wagner. 2017 · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone. 2017 · 2017
Earlier work this paper cites.
CARLA: An open urban driving simulator. In Conference on robot learning . PMLR, 1–16
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. 2017 · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg. 2017 · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning . PMLR, 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Formal Guarantees on the Robustness of a Classifier against Adversarial Manipulation. In NIPS
Matthias Hein and Maksym Andriushchenko. 2017 · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Fairness in reinforcement learning. In International conference on machine learning . PMLR, 1617–1626
Shahin Jabbari, Matthew Joseph, Michael Kearns, Jamie Morgenstern, and Aaron Roth. 2017 · 2017
Earlier work this paper cites.
Delving into adversarial attacks on deep policies
Jernej Kos and Dawn Song. 2017 · 2017
Earlier work this paper cites.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg. 2017 · 2017
Earlier work this paper cites.
Tactics of adversarial attack on deep reinforcement learning agents
Yen-Chen Lin, Zhang-Wei Hong, Yuan-Hong Liao, Meng-Li Shih, Ming-Yu Liu, and Min Sun. 2017a · 2017
Earlier work this paper cites.
Detecting adversarial attacks on neural network policies with visual foresight
Yen-Chen Lin, Ming-Yu Liu, Min Sun, and Jia-Bin Huang. 2017b · 2017
Earlier work this paper cites.
Adversarially robust policy learning: Active construction of physically-plausible perturbations. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 3932–3939
Ajay Mandlekar, Yuke Zhu, Animesh Garg, Li Fei-Fei, and Silvio Savarese. 2017 · 2017
Earlier work this paper cites.
Predicting trust in human control of swarms via inverse reinforcement learning. In 2017 26th ieee international symposium on robot and human interactive communication (ro-man) . IEEE, 528–533
Changjoo Nam, Phillip Walker, Michael Lewis, and Katia Sycara. 2017 · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning. In International Conference on Machine Learning . PMLR, 2817–2826
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. 2017 · 2017
Earlier work this paper cites.
Certifying some distributional robustness with principled adversarial training
Aman Sinha, Hongseok Namkoong, Riccardo Volpi, and John Duchi. 2017 · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS) . IEEE, 23–30
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
A convex optimization approach to distributionally robust Markov decision processes with Wasserstein distance
Insoon Yang. 2017 · 2017
Earlier work this paper cites.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C Karen Liu, and Greg Turk. 2017 · 2017
Earlier work this paper cites.
Learning to attack: Adversarial transformation networks. In Thirty-second aaai conference on artificial intelligence
Shumeet Baluja and Ian Fischer. 2018 · 2018
Earlier work this paper cites.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa. 2018 · 2018
Cited alongside, same era.
On the effectiveness of interval bound propagation for training verifiably robust models
Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli. 2018 · 2018
Cited alongside, same era.
Online robust policy learning in the presence of unknown adversaries
Aaron Havens, Zhanhong Jiang, and Soumik Sarkar. 2018 · 2018
Cited alongside, same era.
Continual reinforcement learning with complex synapses. In International Conference on Machine Learning . PMLR, 2497–2506
Christos Kaplanis, Murray Shanahan, and Claudia Clopath. 2018 · 2018
Cited alongside, same era.
Learning-based model predictive control for safe exploration. In 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 6059–6066
Certified adversarial robustness for deep reinforcement learning. In Conference on Robot Learning . PMLR, 1328–1337
Björn Lütjens, Michael Everett, and Jonathan P How. 2020 · 2020
Later among the works it cites.
Curriculum in gradient-based meta-reinforcement learning
Bhairav Mehta, Tristan Deleu, Sharath Chandra Raparthy, Chris J Pal, and Liam Paull. 2020a · 2020
Later among the works it cites.
Skills for physical artificial intelligence
Aslan Miriyev and Mirko Kovač. 2020 · 2020
Later among the works it cites.
Locally private distributed reinforcement learning
Hajime Ono and Tsubasa Takahashi. 2020 · 2020
Later among the works it cites.
Teacher algorithms for curriculum learning of deep rl in continuously parameterized environments. In Conference on Robot Learning . PMLR, 835–853
Rémy Portelas, Cédric Colas, Katja Hofmann, and Pierre-Yves Oudeyer. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause. 2018 · 2018
Cited alongside, same era.
An Environment for Autonomous Driving Decision-Making
Edouard Leurent. 2018 · 2018
Cited alongside, same era.
The role of trust in human-robot interaction
Michael Lewis, Katia Sycara, and Phillip Walker. 2018 · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano. 2018 · 2018
Cited alongside, same era.
Experimental design for cost-aware learning of causal graphs
Erik Lindgren, Murat Kocaoglu, Alexandros G Dimakis, and Sriram Vishwanath. 2018 · 2018
Cited alongside, same era.
Towards Deep Learning Models Resistant to Adversarial Attacks. In International Conference on Learning Representations
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Anusha Nagabandi, Ignasi Clavera, Simin Liu, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn. 2018a · 2018
Cited alongside, same era.
Deep online learning via meta-learning: Continual adaptation for model-based rl
Anusha Nagabandi, Chelsea Finn, and Sergey Levine. 2018b · 2018
Cited alongside, same era.
Later among the works it cites.
Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning
Younggyo Seo, Kimin Lee, Ignasi Clavera Gilaberte, Thanard Kurutach, Jinwoo Shin, and Pieter Abbeel. 2020 · 2020
Later among the works it cites.
Deep Reinforcement Learning with Robust and Smooth Policy. In International Conference on Machine Learning . PMLR, 8707–8718
Qianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang, and Tuo Zhao. 2020 · 2020
Later among the works it cites.
Learning in Markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff. 2020 · 2020
Later among the works it cites.
Formulazero: Distributionally robust online adaptation via offline population synthesis. In International Conference on Machine Learning . PMLR, 8992–9004
Aman Sinha, Matthew O’Kelly, Hongrui Zheng, Rahul Mangharam, John Duchi, and Russ Tedrake. 2020 · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
Krishnan Srinivasan, Benjamin Eysenbach, Sehoon Ha, Jie Tan, and Chelsea Finn. 2020 · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods. In International Conference on Machine Learning . PMLR, 9133–9143
Adam Stooke, Joshua Achiam, and Pieter Abbeel. 2020 · 2020
Later among the works it cites.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks
Brijen Thananjeyan, Ashwin Balakrishna, Ugo Rosolia, Felix Li, Rowan McAllister, Joseph E Gonzalez, Sergey Levine, Francesco Borrelli, and Ken Goldberg. 2020 · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal. 2020 · 2020
Later among the works it cites.
MDP homomorphic networks: Group symmetries in reinforcement learning
Elise van der Pol, Daniel Worrall, Herke van Hoof, Frans Oliehoek, and Max Welling. 2020 · 2020
Later among the works it cites.
A survey of multi-task deep reinforcement learning
Nelson Vithayathil Varghese and Qusay H Mahmoud. 2020 · 2020
Later among the works it cites.
Reinforcement learning with perturbed rewards. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 6202–6209
Jingkang Wang, Yang Liu, and Bo Li. 2020 · 2020
Later among the works it cites.
Task-agnostic online reinforcement learning with an infinite mixture of gaussian processes
Mengdi Xu, Wenhao Ding, Jiacheng Zhu, Zuxin Liu, Baiming Chen, and Ding Zhao. 2020 · 2020
Later among the works it cites.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge. 2020 · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine. 2020d · 2020
Later among the works it cites.
Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 21024–21037
Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan Liu, Duane Boning, and Cho-Jui Hsieh. 2020a · 2020
Later among the works it cites.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross. 2020e · 2020
Later among the works it cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey. In 2020 IEEE Symposium Series on Computational Intelligence (SSCI) . IEEE, 737–744
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund. 2020 · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín. 2020 · 2020
Later among the works it cites.
Deep probabilistic accelerated evaluation: A robust certifiable rare-event simulation methodology for black-box safety-critical systems. In International Conference on Artificial Intelligence and Statistics . PMLR, 595–603
Mansur Arief, Zhiyuan Huang, Guru Koushik Senthil Kumar, Yuanlu Bai, Shengyi He, Wenhao Ding, Henry Lam, and Ding Zhao. 2021 · 2021
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora and Prashant Doshi. 2021 · 2021
Later among the works it cites.
Invariant Causal Imitation Learning for Generalizable Policies
Ioana Bica, Daniel Jarrett, and Mihaela van der Schaar. 2021 · 2021
Later among the works it cites.
Context-Aware Safe Reinforcement Learning for Non-Stationary Environments
Baiming Chen, Zuxin Liu, Jiacheng Zhu, Mengdi Xu, Wenhao Ding, and Ding Zhao. 2021 · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization. In International Conference on Artificial Intelligence and Statistics . PMLR, 3304–3312
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic. 2021 · 2021
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester. 2021 · 2021
Later among the works it cites.
Certifiable Robustness to Adversarial State Uncertainty in Deep Reinforcement Learning
Michael Everett, Björn Lütjens, and Jonathan P How. 2021 · 2021
Later among the works it cites.
Explaining a Deep Reinforcement Learning Docking Agent Using Linear Model Trees with User Adapted Visualization
Vilde B Gjærum, Inga Strümke, Ole Andreas Alsos, and Anastasios M Lekkas. 2021 · 2021
Later among the works it cites.
Embodied intelligence via learning and evolution
Agrim Gupta, Silvio Savarese, Surya Ganguli, and Li Fei-Fei. 2021 · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation. In 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 13611–13617
Nicklas Hansen and Xiaolong Wang. 2021 · 2021
Later among the works it cites.
What would jiminy cricket do? towards agents that behave morally
Dan Hendrycks, Mantas Mazeika, Andy Zou, Sahil Patel, Christine Zhu, Jesus Navarro, Dawn Song, Bo Li, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Inaam Ilahi, Muhammad Usama, Junaid Qadir, Muhammad Umar Janjua, Ala Al-Fuqaha, Dinh Thai Hoang, and Dusit Niyato. 2021 · 2021
Later among the works it cites.
Prioritized level replay. In International Conference on Machine Learning . PMLR, 4940–4950
Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel. 2021 · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel. 2021 · 2021
Later among the works it cites.
Policy smoothing for provably robust reinforcement learning
Aounon Kumar, Alexander Levine, and Soheil Feizi. 2021 · 2021
Later among the works it cites.
Recurrent neural network controllers for signal temporal logic specifications subject to safety constraints
Wenliang Liu, Noushin Mehdipour, and Calin Belta. 2021 · 2021
Later among the works it cites.
Distributionally Robust Partially Observable Markov Decision Process with Moment-Based Ambiguity
Hideaki Nakao, Ruiwei Jiang, and Siqian Shen. 2021 · 2021
Later among the works it cites.
Robust Deep Reinforcement Learning through Adversarial Loss. In Advances in Neural Information Processing Systems , A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (Eds.)
Tuomas Oikarinen, Wang Zhang, Alexandre Megretski, Luca Daniel, and Tsui-Wei Weng. 2021 · 2021
Later among the works it cites.
Deep reinforcement learning in production systems: a systematic literature review
Marcel Panzer and Benedict Bender. 2022 · 2021
Later among the works it cites.
Learning from Demonstrations Using Signal Temporal Logic in Stochastic and Continuous Domains
Aniruddh Gopinath Puranic, Jyotirmoy Deshmukh, and Stefanos Nikolaidis. 2021 · 2021
Later among the works it cites.
Learning neural causal models with active interventions
Nino Scherrer, Olexa Bilaniuk, Yashas Annadani, Anirudh Goyal, Patrick Schwab, Bernhard Schölkopf, Michael C Mozer, Yoshua Bengio, Stefan Bauer, and Nan Rosemary Ke. 2021 · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg. 2021 · 2021
Later among the works it cites.
Explainable ai and reinforcement learning—a systematic review of current approaches and trends
Lindsay Wells and Tomasz Bednarz. 2021 · 2021
Later among the works it cites.
arXiv preprint arXiv:2106.09292 (2021)
Fan Wu, Linyi Li, Zijian Huang, Yevgeniy Vorobeychik, Ding Zhao, and Bo Li. 2021 · 2021
Later among the works it cites.
Accelerated Policy Evaluation: Learning Adversarial Environments with Adaptive Importance Sampling
Mengdi Xu, Peide Huang, Fengpei Li, Jiacheng Zhu, Xuewei Qi, Kentaro Oguchi, Zhiyuan Huang, Henry Lam, and Ding Zhao. 2021 · 2021
Later among the works it cites.
Phy-q: A benchmark for physical reasoning
Cheng Xue, Vimukthini Pinto, Chathura Gamage, Ekaterina Nikonova, Peng Zhang, and Jochen Renz. 2021 · 2021
Later among the works it cites.
Reinforcement Learning in Healthcare: A Survey
Chao Yu, Jiming Liu, Shamim Nemati, and Guosheng Yin. 2021 · 2021
Later among the works it cites.
Transform2Act: Learning a Transform-and-Control Policy for Efficient Agent Design
Ye Yuan, Yuda Song, Zhengyi Luo, Wen Sun, and Kris Kitani. 2021b · 2021
Later among the works it cites.
Zhaocong Yuan, Adam W Hall, Siqi Zhou, Lukas Brunke, Melissa Greeff, Jacopo Panerati, and Angela P Schoellig. 2021a · 2021
Later among the works it cites.
Robust Reinforcement Learning on State Observations with Learned Optimal Adversary. In International Conference on Learning Representations
Huan Zhang, Hongge Chen, Duane S Boning, and Cho-Jui Hsieh. 2021 · 2021
Later among the works it cites.
Constrained Policy Optimization via Bayesian World Models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause. 2022 · 2022
Closest in time.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig. 2022 · 2022
Closest in time.
Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal Reasoning
Wenhao Ding, Haohong Lin, Bo Li, and Ding Zhao. 2022a · 2022
Closest in time.
A Survey on Safety-critical Scenario Generation from Methodological Perspective
Wenhao Ding, Chejian Xu, Haohong Lin, Bo Li, and Ding Zhao. 2022b · 2022
Closest in time.
Robust Markov Decision Processes: Beyond Rectangularity
Vineet Goyal and Julien Grand-Clement. 2022 · 2022
Closest in time.
BULLET-SAFETY-GYM: AFRAMEWORK FOR CONSTRAINED REINFORCEMENT LEARNING
Sven Gronauer. 2022 · 2022
Closest in time.
A Review of Safe Reinforcement Learning: Methods, Theory and Applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll. 2022 · 2022
Closest in time.
Robust Reinforcement Learning as a Stackelberg Game via Adaptively-Regularized Adversarial Training
Peide Huang, Mengdi Xu, Fei Fang, and Ding Zhao. 2022 · 2022
Closest in time.
On the Robustness of Safe Reinforcement Learning under Observational Perturbations
Zuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang, Jie Tan, Bo Li, and Ding Zhao. 2022b · 2022
Closest in time.
CompoSuite: A Compositional Reinforcement Learning Benchmark
Jorge A Mendez, Marcel Hussing, Meghna Gummadi, and Eric Eaton. 2022 · 2022
Closest in time.
Robust Reinforcement Learning: A Review of Foundations and Recent Advances
Janosch Moos, Kay Hansel, Hany Abdulsamad, Svenja Stark, Debora Clever, and Jan Peters. 2022 · 2022
Closest in time.
Safe driving via expert guided policy optimization. In Conference on Robot Learning . PMLR, 1554–1563
Zhenghao Peng, Quanyi Li, Chunxiao Liu, and Bolei Zhou. 2022 · 2022
Closest in time.
Sauté RL: Almost surely safe reinforcement learning using state augmentation. In International Conference on Machine Learning . PMLR, 20423–20443
Aivar Sootla, Alexander I Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David H Mguni, Jun Wang, and Haitham Ammar. 2022 · 2022
Closest in time.
Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RL. In International Conference on Learning Representations
Yanchao Sun, Ruijie Zheng, Yongyuan Liang, and Furong Huang. 2022 · 2022
Closest in time.
COPA: Certifying Robust Policies for Offline Reinforcement Learning against Poisoning Attacks
Fan Wu, Linyi Li, Chejian Xu, Huan Zhang, Bhavya Kailkhura, Krishnaram Kenthapadi, Ding Zhao, and Bo Li. 2022b · 2022
Closest in time.
SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles
Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Shuai Wang, Yihan He, Hanjiang Hu, Ding Zhao, and Bo Li. 2022 · 2022
Closest in time.
Towards Safe Reinforcement Learning with a Safety Editor Policy
Haonan Yu, Wei Xu, and Haichao Zhang. 2022 · 2022
Closest in time.
Robust Deep Reinforcement Learning with adversarial attacks. In 17th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2018 . International Foundation for Autonomous Agents and Multiagent Systems (IFAAMAS), 2040–2042
Anay Pattanaik, Zhenyi Tang, Shuijing Liu, Gautham Bommannan, and Girish Chowdhary. 2018 · 2042
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning. In International conference on machine learning . PMLR, 2048–2056
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman. 2020 · 2056
Closest in time.