Fetching the paper…
Reading the bibliography…
As the complexity of tasks addressed through reinforcement learning (RL) increases, the definition of reward functions also has become highly complicated.
E. Altman, Constrained Markov decision processes . CRC Press, 1999, vol. 7
1999
Earlier work this paper cites.
K. Van Moffaert, M. M. Drugan, and A. Nowé, “Scalarized multi-objective reinforcement learning: Novel design techniques,” in Proceedings of IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning , 2013
2013
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of International Conference on Machine Learning , 2015
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proceedings of International Conference on Machine Learning , 2017
2017
Earlier work this paper cites.
X. B. Peng, P. Abbeel, S. Levine, and M. Van de Panne, “DeepMimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Transactions On Graphics (TOG) , vol. 37, no. 4, pp. 1–14, 2018
2018
Earlier work this paper cites.
R. Yang, X. Sun, and K. Narasimhan, “A generalized algorithm for multi-objective reinforcement learning and policy adaptation,” Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning quadrupedal locomotion over challenging terrain,” Science Robotics , vol. 5, no. 47, 2020
2020
Earlier work this paper cites.
J. Xu, Y. Tian, P. Ma, D. Rus, S. Sueda, and W. Matusik, “Prediction-guided multi-objective reinforcement learning for continuous robot control,” in Proceedings of International Conference on Machine Learning , 2020
2020
Earlier work this paper cites.
A. Abdolmaleki, S. Huang, L. Hasenclever, M. Neunert, F. Song, M. Zambelli, M. Martins, N. Heess, R. Hadsell, and M. Riedmiller, “A distributional view on multi-objective policy optimization,” in Proceedings of International Conference on Machine Learning , 2020
2020
Earlier work this paper cites.
A. Stooke, J. Achiam, and P. Abbeel, “Responsive safety in reinforcement learning by pid lagrangian methods,” in Proceedings of International Conference on Machine Learning , 2020
2020
Earlier work this paper cites.
A. Kumar, Z. Fu, D. Pathak, and J. Malik, “RMA: Rapid motor adaptation for legged robots,” in Robotics: Science and Systems , 2021
2021
Earlier work this paper cites.
J. Siekmann, K. Green, J. Warila, A. Fern, and J. Hurst, “Blind bipedal stair traversal via sim-to-real reinforcement learning,” in Robotics: Science and Systems , 2021
2021
Cited alongside, same era.
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al. , “Isaac gym: High performance gpu based physics simulation for robot learning,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
L. Smith, J. C. Kew, X. B. Peng, S. Ha, J. Tan, and S. Levine, “Legged robots that keep on learning: Fine-tuning locomotion policies in the real world,” in Proceedings of International Conference on Robotics and Automation , 2022
2022
Cited alongside, same era.
G. Ji, J. Mun, H. Kim, and J. Hwangbo, “Concurrent training of a control policy and a state estimator for dynamic and robust legged locomotion,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 4630–4637, 2022
2022
Z. Li, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath, “Robust and versatile bipedal jumping control through reinforcement learning,” in Robotics: Science and Systems , 2023
2023
Later among the works it cites.
T. Basaklar, S. Gumussoy, and U. Ogras, “PD-MORL: Preference-driven multi-objective reinforcement learning algorithm,” in Proceedings of International Conference on Learning Representations , 2023
2023
Later among the works it cites.
X.-Q. Cai, P. Zhang, L. Zhao, J. Bian, M. Sugiyama, and A. Llorens, “Distributional pareto-optimal multi-objective reinforcement learning,” in Advances in Neural Information Processing Systems , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
X. B. Peng, Y. Guo, L. Halper, S. Levine, and S. Fidler, “ASE: Large-scale reusable adversarial skill embeddings for physically simulated characters,” ACM Transactions On Graphics (TOG) , vol. 41, no. 4, pp. 1–17, 2022
2022
Cited alongside, same era.
L. Zhang, L. Shen, L. Yang, S.-Y. Chen, B. Yuan, X. Wang, and D. Tao, “Penalized proximal policy optimization for safe reinforcement learning,” in Proceedings of International Joint Conference on Artificial Intelligence , 2022
2022
Cited alongside, same era.
P. Kyriakis and J. Deshmukh, “Pareto policy adaptation,” in Proceedings of International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
C. F. Hayes, R. Rădulescu, E. Bargiacchi, J. Källström, M. Macfarlane, M. Reymond, T. Verstraeten, L. M. Zintgraf, R. Dazeley, F. Heintz, et al. , “A practical guide to multi-objective reinforcement learning and planning,” Autonomous Agents and Multi-Agent Systems , vol. 36, no. 1, p. 26, 2022
2022
Cited alongside, same era.
S. Huang, A. Abdolmaleki, G. Vezzani, P. Brakel, D. J. Mankowitz, M. Neunert, S. Bohez, Y. Tassa, N. Heess, M. Riedmiller, et al. , “A constrained multi-objective reinforcement learning framework,” in Proceedings of Conference on Robot Learning , 2022
2022
Cited alongside, same era.
G. B. Margolis and P. Agrawal, “Walk these ways: Tuning robot control for generalization with multiplicity of behavior,” in Proceedings of Conference on Robot Learning , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Lu, D. Herman, and Y. Yu, “Multi-objective reinforcement learning: Convexity, stationarity and Pareto optimality,” in Proceedings of International Conference on Learning Representations , 2023
2023
Later among the works it cites.
A. Agarwal, A. Kumar, J. Malik, and D. Pathak, “Legged locomotion in challenging terrains using egocentric vision,” in Proceedings of Conference on robot learning , 2023
2023
Later among the works it cites.
X. Cheng, K. Shi, A. Agarwal, and D. Pathak, “Extreme parkour with legged robots,” in Proceedings of International Conference on Robotics and Automation , 2024
2024
Closest in time.
R. Grandia, E. Knoop, M. A. Hopkins, G. Wiedebach, J. Bishop, S. Pickles, D. Müller, and M. Bächer, “Design and control of a bipedal robotic character,” in Robotics: Science and Systems , 2024
2024
Closest in time.
2024
Closest in time.
Y. Kim, H. Oh, J. Lee, J. Choi, G. Ji, M. Jung, D. Youm, and J. Hwangbo, “Not only rewards but also constraints: Applications on legged robot locomotion,” IEEE Transactions on Robotics , 2024
2024
Closest in time.
U. Robotics, “unitree_ros,” https://github.com/unitreerobotics/unitree_ros , 2024, gitHub repository
2024
Closest in time.