Fetching the paper…
Reading the bibliography…
Training and deploying reinforcement learning (RL) policies for robots, especially in accomplishing specific tasks, presents substantial challenges.
K. Åström and P. Eykhoff, “System identification-A survey,”
1971
Earlier work this paper cites.
M. Quigley, K. Conley, B. Gerkey, J. Faust, T. Foote, J. Leibs, R. Wheeler, A. Y. Ng
2009
Earlier work this paper cites.
M. Kalakrishnan, J. Buchli, P. Pastor, M. Mistry, and S. Schaal, “Learning, planning, and control for quadruped locomotion over challenging terrain,” in
2011
Earlier work this paper cites.
M. P. Deisenroth and C. E. Rasmussen, “PILCO: A model-based and data-efficient approach to policy search,”
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,”
2012
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne, “DeepMimic: Example-guided deep reinforcement learning of physics-based character skills,” in
2018
Earlier work this paper cites.
T. Apgar, P. Clary, K. Green, A. Fern, and J. W. Hurst, “Fast Online Trajectory Optimization for the Bipedal Robot Cassie.” in
2018
Earlier work this paper cites.
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Hwangbo, J. Lee, L. Wellhausen, H. Kolvenbach, and M. Hutter, “Learning Agile and Dynamic Motor Skills for Legged Robots,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. B. Slaoui, W. R. Clements, J. N. Foerster, and S. Toth, “Robust Domain Randomization for Reinforcement Learning,” 2020. [Online]. Available:
2020
Earlier work this paper cites.
J. Lee, J. Hwangbo, L. Wellhausen, H. Kolvenbach, and M. Hutter, “Learning quadrupedal locomotion over challenging terrain,”
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
2020
Earlier work this paper cites.
Salvato, Erica and Fenu, Gianfranco and Medvet, Eric and Pellegrino, Felice Andrea, “Crossing the Reality Gap: A Survey on Sim-to-Real Transferability of Robot Controllers in Reinforcement Learning,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Duan, J. Dao, K. Green, T. Apgar, A. Fern, and J. Hurst, “Learning Task Space Actions for Bipedal Locomotion,” in
2021
Earlier work this paper cites.
J. Siekmann, Y. Godse, A. Fern, and J. Hurst, “Sim-to-real learning of all common bipedal gaits via periodic reward composition,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Wang and H. Li, “Code completion by modeling flattened abstract syntax trees as graphs,” in
2021
Earlier work this paper cites.
N. Jiang, T. Lutellier, and L. Tan, “Cure: Code-aware neural machine translation for automatic program repair,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Acosta, W. Yang, and M. Posa, “Validating Robotics Simulators on Real-World Impacts,”
2022
Earlier work this paper cites.
E. Šutinys, U. Samukaitė-Bubnienė, and V. Bučinskas, “Advanced Applications of Industrial Robotics: New Trends and Possibilities,”
2022
Earlier work this paper cites.
F. Ciardo and A. Wykowska, “Humanoid robot passes for human in joint task experiment,” in
2022
Cited alongside, same era.
R. Toro Icarte, T. Q. Klassen, R. Valenzano, and S. A. McIlraith, “Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning,”
2022
Cited alongside, same era.
L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning,”
2022
Cited alongside, same era.
T. Korbak, H. Elsahar, G. Kruszewski, and M. Dymetman, “On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting,” in
2022
Cited alongside, same era.
N. Rudin, D. Hoeller, P. Reist, and M. Hutter, “Learning to walk in minutes using massively parallel deep reinforcement learning,” in
2022
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar, “Eureka: Human-level reward design via coding large language models,”
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Green, Kevin and Warila, John and Hatton, Ross L. and Hurst, Jonathan, “Motion planning for agile legged locomotion using failure margin constraints,” in
2022
Cited alongside, same era.
W. Huang, N. Gopalan, M. Ahn
2022
Cited alongside, same era.
W. Huang, A. Brohan, N. Gopalan
2022
Cited alongside, same era.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar,
2022
Cited alongside, same era.
A. Brunnbauer
2022
Cited alongside, same era.
F. Muratore, F. Ramos, G. Turk, W. Yu, M. Gienger, and J. Peters, “Robot learning from randomized simulations: A review,”
2022
Cited alongside, same era.
H. Duan, A. Malik, J. Dao, A. Saxena, K. Green, J. Siekmann, A. Fern, and J. Hurst, “Sim-to-Real Learning of Footstep-Constrained Bipedal Dynamic Walking,” in
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
F. Zeng, W. Gan, Y. Wang, N. Liu, and P. S. Yu, “Large language models for robotics: A survey,”
2023
Later among the works it cites.
J. Betz and H. Zheng, “Bypassing the Simulation-to-Reality Gap: Online Reinforcement Learning Using a Supervisor,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou
2023
Later among the works it cites.
2023
Later among the works it cites.
P. Katara, Z. Xian, and K. Fragkiadaki, “Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models,” in
2024
Closest in time.
F. Wu, Z. Gu, H. Wu, A. Wu, and Y. Zhao, “Infer and Adapt: Bipedal Locomotion Reward Learning from Demonstrations via Inverse Reinforcement Learning,” in
2024
Closest in time.
2024
Closest in time.
D. Ernst and A. Louette, “Introduction to reinforcement learning,”
2024
Closest in time.
2024
Closest in time.
Y. J. Ma, W. Liang, H. Wang, S. Wang, Y. Zhu, L. Fan, O. Bastani, and D. Jayaraman, “DrEureka: Language Model Guided Sim-To-Real Transfer,” in
2024
Closest in time.
F. Wu, Z. Gu, H. Wu, A. Wu, and Y. Zhao, “Infer and adapt: Bipedal locomotion reward learning from demonstrations via inverse reinforcement learning,” in
2024
Closest in time.
T. Haarnoja, B. Moran, G. Lever, S. H. Huang, D. Tirumala, J. Humplik, M. Wulfmeier, S. Tunyasuvunakool, N. Y. Siegel, R. Hafner
2024
Closest in time.
S. H. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor, “Chatgpt for robotics: Design principles and model abilities,”
2024
Closest in time.
Y. Kim, H. Oh, J. Lee, J. Choi, G. Ji, M. Jung, D. Youm, and J. Hwangbo, “Not only rewards but also constraints: Applications on legged robot locomotion,”
2024
Closest in time.
Anthropic, “Claude 3.5 sonnet,” 2024, accessed: 2025-02-21. [Online]. Available:
2025
Closest in time.
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi
2025
Closest in time.
W. Xia, D. Wang, X. Pang, Z. Wang, B. Zhao, D. Hu, and X. Li, “Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMs,” in
2080
Closest in time.