Fetching the paper…
Reading the bibliography…
Human intervention is an effective way to inject human knowledge into the training loop of reinforcement learning, which can bring fast learning and ensured training safety.
Pattern recognition and adaptive control
Bernard Widrow · 1964
Earlier work this paper cites.
Behavioral cloning a correction
Rui Camacho and Donald Michie · 1995
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
The panda3d graphics engine
Mike Goslin and Mark R Mine · 2004
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Infinite time horizon maximum causal entropy inverse reinforcement learning
Michael Bloem and Nicholas Bambos · 2014
Earlier work this paper cites.
Policy shaping with human teachers
Thomas Cederborg, Ishaan Grover, Charles L Isbell, and Andrea L Thomaz · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
Jiakai Zhang and Kyunghyun Cho · 2016
Earlier work this paper cites.
Agent-agnostic human-in-the-loop reinforcement learning
David Abel, John Salvatier, Andreas Stuhlmüller, and Owain Evans · 2017
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
CARLA: an open urban driving simulator
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López, and Vladlen Koltun · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Cited alongside, same era.
Where to add actions in human-in-the-loop reinforcement learning
Travis Mandel, Yun-En Liu, Emma Brunskill, and Zoran Popović · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Cited alongside, same era.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
Michael Kelly, Chelsea Sidrane, Katherine Driggs-Campbell, and Mykel J Kochenderfer · 2019
Later among the works it cites.
Learning reward functions by integrating human demonstrations and preferences
Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh · 2019
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Active reward learning from critiques
Yuchen Cui and Scott Niekum · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, and Jan Peters · 2018
Cited alongside, same era.
Homanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine, Florian Shkurti, and Animesh Garg · 2020
Later among the works it cites.
Learning to walk in the real world with minimal human effort, 2020
Sehoon Ha, Peng Xu, Zhenyu Tan, Sergey Levine, and Jie Tan · 2020
Later among the works it cites.
Learning a decision module by imitating driver’s control behaviors
Junning Huang, Sirui Xie, Jiankai Sun, Qiurui Ma, Chunxiao Liu, Dahua Lin, and Bolei Zhou · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Human-in-the-loop imitation learning using remote teleoperation
Ajay Mandlekar, Danfei Xu, Roberto Martín-Martín, Yuke Zhu, Li Fei-Fei, and Silvio Savarese · 2020
Later among the works it cites.
Learning from interventions
Jonathan Spencer, Sanjiban Choudhury, Matthew Barnes, Matthew Schmittle, Mung Chiang, Peter Ramadge, and Siddhartha Srinivasa · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Neuro-symbolic program search for autonomous driving decision module design
Jiankai Sun, Hao Sun, Tian Han, and Bolei Zhou · 2020
Later among the works it cites.
Continual learning of control primitives: Skill discovery via reset-games
Kelvin Xu, Siddharth Verma, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
Lin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang, and Subbarao Kambhampati · 2021
Later among the works it cites.
Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning, 2021
Ryan Hoque, Ashwin Balakrishna, Ellen Novoseller, Albert Wilcox, Daniel S. Brown, and Ken Goldberg · 2021
Later among the works it cites.
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Quanyi Li, Zhenghao Peng, Zhenghai Xue, Qihang Zhang, and Bolei Zhou · 2021
Later among the works it cites.
Safe driving via expert guided policy optimization
Zhenghao Peng, Quanyi Li, Chunxiao Liu, and Bolei Zhou · 2021
Later among the works it cites.
Human-in-the-loop deep reinforcement learning with application to autonomous driving
Jingda Wu, Zhiyu Huang, Chao Huang, Zhongxu Hu, Peng Hang, Yang Xing, and Chen Lv · 2021
Later among the works it cites.
Recent advances in leveraging human guidance for sequential decision-making tasks
Ruohan Zhang, Faraz Torabi, Garrett Warnell, and Peter Stone · 2021
Later among the works it cites.