Proximal policy optimization algorithms
Original
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
A Survey of Preference-Based Reinforcement Learning Methods
Wirth, C.; Akrour, R.; Neumann, G.; and Fürnkranz, J. 2017 · 2017
Cited alongside, same era.
Motion planning among dynamic, decision-making agents with deep reinforcement learning
Everett, M.; Chen, Y. F.; and How, J. P. 2018 · 2018
Cited alongside, same era.
QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
Original
Kalashnikov, D.; Irpan, A.; Pastor, P.; Ibarz, J.; Herzog, A.; Jang, E.; Quillen, D.; Holly, E.; Kalakrishnan, M.; Vanhoucke, V.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A.; McGrew, B.; Andrychowicz, M.; Zaremba, W.; and Abbeel, P. 2018 · 2018
Cited alongside, same era.
Deepmind control suite
Original
Tassa, Y.; Doron, Y.; Muldal, A.; Erez, T.; Li, Y.; Casas, D. d. L.; Budden, D.; Abdolmaleki, A.; Merel, J.; Lefrancq, A.; et al. 2018 · 2018
Cited alongside, same era.
Robot navigation in crowded environments using deep reinforcement learning
Liu, L.; Dugas, D.; Cesari, G.; Siegwart, R.; and Dubé, R. 2020 · 2020
Cited alongside, same era.
Reinforcement learning for robotic manipulation using simulated locomotion demonstrations
Kilinc, O.; and Montana, G. 2021 · 2021
Cited alongside, same era.
PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Original
Lee, K.; Smith, L.; and Abbeel, P. 2021 · 2021
Cited alongside, same era.
Reinforcement Learning in Sparse-Reward Environments With Hindsight Policy Gradients
Rauber, P.; Ummadisingu, A.; Mutz, F.; and Schmidhuber, J. 2021 · 2021
Cited alongside, same era.
Technical Report for ”Dealing with Sparse Rewards in Continuous Control Robotics via Heavy-Tailed Policy Optimization”
2022 · 2022
Cited alongside, same era.
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning
Original
Chakraborty, S.; Bedi, A. S.; Koppel, A.; Manocha, D.; Wang, H.; Wang, M.; and Huang, F. 2023a
Cited in the paper.