Fetching the paper…
Reading the bibliography…
When learning common skills like driving, beginners usually have domain experts standing by to ensure the safety of the learning process.
Pattern recognition and adaptive control
B. Widrow · 1964
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
The panda3d graphics engine
M. Goslin and M. R. Mine · 2004
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
J. Zhang and K. Cho · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
Where to add actions in human-in-the-loop reinforcement learning
T. Mandel, Y.-E. Liu, E. Brunskill, and Z. Popović · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Agent-agnostic human-in-the-loop reinforcement learning
D. Abel, J. Salvatier, A. Stuhlmüller, and O. Evans · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
A lyapunov-based approach to safe reinforcement learning
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh · 2018
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
W. Saunders, G. Sastry, A. Stuhlmueller, and O. Evans · 2018
Earlier work this paper cites.
Safe exploration in continuous action spaces
G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica · 2018
Cited alongside, same era.
Conservative safety critics for exploration
H. Bharadhwaj, A. Kumar, N. Rhinehart, S. Levine, F. Shkurti, and A. Garg · 2020
Later among the works it cites.
Explanation augmented feedback in human-in-the-loop reinforcement learning
L. Guan, M. Verma, S. Guo, R. Zhang, and S. Kambhampati · 2020
Later among the works it cites.
Balancing constraints and rewards with meta-gradient d4pg
D. A. Calian, D. J. Mankowitz, T. Zahavy, Z. Xu, J. Oh, N. Levine, and T. Mann · 2020
Later among the works it cites.
Learning for safety-critical control with control barrier functions
A. Taylor, A. Singletary, Y. Yue, and A. Ames · 2020
Later among the works it cites.
Ipo: Interior-point policy optimization under constraints
Y. Liu, J. Ding, and X. Liu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-dimensional continuous control using generalized advantage estimation, 2018
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2018
Cited alongside, same era.
Learning to drive in a day
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah · 2019
Cited alongside, same era.
A review of reinforcement learning for autonomous building energy management
K. Mason and S. Grijalva · 2019
Cited alongside, same era.
Open-sourced reinforcement learning environments for surgical robotics
F. Richter, R. K. Orosco, and M. C. Yip · 2019
Cited alongside, same era.
Sqil: Imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2019
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
Lyapunov barrier policy optimization
H. Sikchi, W. Zhou, and D. Held · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
K. Srinivasan, B. Eysenbach, S. Ha, J. Tan, and C. Finn · 2020
Later among the works it cites.
Neuro-symbolic program search for autonomous driving decision module design
J. Sun, H. Sun, T. Han, and B. Zhou · 2020
Later among the works it cites.
Learning a decision module by imitating driver’s control behaviors
J. Huang, S. Xie, J. Sun, Q. Ma, C. Liu, D. Lin, and B. Zhou · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Learning from interventions
J. Spencer, S. Choudhury, M. Barnes, M. Schmittle, M. Chiang, P. Ramadge, and S. Srinivasa · 2020
Later among the works it cites.
Learning to walk in the real world with minimal human effort, 2020
S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan · 2020
Later among the works it cites.
Adversarial inverse reinforcement learning with self-attention dynamics model
J. Sun, L. Yu, P. Dong, B. L, and B. Zhou · 2021
Closest in time.
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Q. Li, Z. Peng, Z. Xue, Q. Zhang, and B. Zhou · 2021
Closest in time.