Fetching the paper…
Reading the bibliography…
A long-term goal of reinforcement learning is to design agents that can autonomously interact and learn in the world.
Don’t do things you can’t undo: reversibility models for generating safe behaviours
Maarja Kruusmaa, Yuri Gavshin, and Adam Eppendahl · 2007
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Preference-based policy learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: A preliminary survey
Christian Wirth and Johannes Fürnkranz · 2013
Earlier work this paper cites.
Learning compound multi-step controllers under unknown dynamics
Weiqiao Han, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
Jiakai Zhang and Kyunghyun Cho · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Earlier work this paper cites.
Deep predictive policy training using reinforcement learning
Ali Ghadirzadeh, Atsuto Maki, Danica Kragic, and Mårten Björkman · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions
Erdem Biyik and Dorsa Sadigh · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Penalizing side effects using stepwise relative reachability
Learning the arrow of time for problems in reinforcement learning
Nasim Rahaman, Steffen Wolf, Anirudh Goyal, Roman Remme, and Yoshua Bengio · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
Krishnan Srinivasan, Benjamin Eysenbach, Sehoon Ha, Jie Tan, and Chelsea Finn · 2020
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal · 2020
Later among the works it cites.
Continual learning of control primitives: Skill discovery via reset-games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Victoria Krakovna, Laurent Orseau, Ramana Kumar, Miljan Martic, and Shane Legg · 2018
Cited alongside, same era.
Episodic curiosity through reachability
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Zhaodong Wang and Matthew E Taylor · 2018
Cited alongside, same era.
Ensembledagger: A bayesian approach to safe imitation learning
Kunal Menda, Katherine Driggs-Campbell, and Mykel J Kochenderfer · 2019
Cited alongside, same era.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
Henry Zhu, Abhishek Gupta, Aravind Rajeswaran, Sergey Levine, and Vikash Kumar · 2019
Cited alongside, same era.
A survey on interactive reinforcement learning: design principles and open challenges
Christian Arzate Cruz and Takeo Igarashi · 2020
Cited alongside, same era.
Kelvin Xu, Siddharth Verma, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
The ingredients of real-world robotic reinforcement learning
Henry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah, Kristian Hartikainen, Avi Singh, Vikash Kumar, and Sergey Levine · 2020
Later among the works it cites.
Safe reinforcement learning via statistical model predictive shielding
Osbert Bastani, Shuo Li, and Anton Xu · 2021
Later among the works it cites.
There is no turning back: A self-supervised approach for reversibility-aware reinforcement learning
Nathan Grinsztajn, Johan Ferret, Olivier Pietquin, Matthieu Geist, et al · 2021
Later among the works it cites.
Abhishek Gupta, Justin Yu, Tony Zhao, Vikash Kumar, Aaron Rovinsky, Kelvin Xu, Thomas Devlin, and Sergey Levine · 2021
Later among the works it cites.
Lazydagger: Reducing context switching in interactive imitation learning
Ryan Hoque, Ashwin Balakrishna, Carl Putterman, Michael Luo, Daniel S Brown, Daniel Seita, Brijen Thananjeyan, Ellen Novoseller, and Ken Goldberg · 2021
Later among the works it cites.
Kimin Lee, Laura Smith, and Pieter Abbeel · 2021
Later among the works it cites.
Safe reinforcement learning by imagining the near future
Garrett Thomas, Yuping Luo, and Tengyu Ma · 2021
Later among the works it cites.
Safe reinforcement learning using advantage-based intervention
Nolan C Wagener, Byron Boots, and Ching-An Cheng · 2021
Later among the works it cites.
Safe continuous control with constrained model-based policy optimization
Moritz A Zanger, Karam Daaboul, and J Marius Zöllner · 2021
Later among the works it cites.
A state-distribution matching approach to non-episodic reinforcement learning
Archit Sharma, Rehaan Ahmad, and Chelsea Finn · 2022
Closest in time.
Skill preferences: Learning to extract and execute robotic skills from human feedback
Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi, Pieter Abbeel, and Michael Laskin · 2022
Closest in time.