Fetching the paper…
Reading the bibliography…
In robot manipulation, Reinforcement Learning (RL) often suffers from low sample efficiency and uncertain convergence, especially in large observation and action spaces.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning (ICML) . PMLR, 2018
2018
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani et al. , “Transporter networks: Rearranging the visual world for robotic manipulation,” in Conference on Robot Learning (CoRL) . PMLR, 2021
2021
Earlier work this paper cites.
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” The Journal of Machine Learning Research , vol. 22, no. 1, pp. 12 348–12 355, 2021
2021
Earlier work this paper cites.
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” in International Conference on Machine Learning (ICML) . PMLR, 2022
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Earlier work this paper cites.
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” in International Conference on Learning Representations (ICLR) , 2022
2022
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “CLIPort: What and where pathways for robotic manipulation,” in Conference on Robot Learning (CoRL) . PMLR, 2022
2022
Earlier work this paper cites.
N. Di Palo, A. Byravan, L. Hasenclever, M. Wulfmeier, N. Heess, and M. Riedmiller, “Towards a unified agent with foundation models,” in Workshop on Reincarnating Reinforcement Learning at ICLR , 2023
2023
Cited alongside, same era.
OpenAI, “GPT-4 technical report,” 2023
2023
Cited alongside, same era.
A. Zeng, M. Attarian, B. Ichter, K. M. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. Florence, “Socratic models: Composing zero-shot multimodal reasoning with language,” in International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “VoxPoser: Composable 3D value maps for robotic manipulation with language models,” in Conference on Robot Learning (CoRL) . PMLR, 2023
2023
Cited alongside, same era.
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K.-H. Lee, M. G. Arenas, H.-T. L. Chiang, T. Erez, L. Hasenclever, J. Humplik, B. Ichter, T. Xiao, P. Xu, A. Zeng, T. Zhang, N. Heess, D. Sadigh, J. Tan, Y. Tassa, and F. Xia, “Language to rewards for robotic skill synthesis,” in Conference on Robot Learning (CoRL) . PMLR, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Du, O. Watkins, Z. Wang, C. Colas, T. Darrell, P. Abbeel, A. Gupta, and J. Andreas, “Guiding pretraining in reinforcement learning with large language models,” in International Conference on Machine Learning (ICML) . PMLR, 2023
2023
Later among the works it cites.
S. Kambhampati, K. Valmeekam, L. Guan, M. Verma, K. Stechly, S. Bhambri, L. P. Saldyt, and A. B. Murthy, “Position: LLMs can’t plan, but can help planning in LLM-modulo frameworks,” in International Conference on Machine Learning (ICML) , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Carta, C. Romac, T. Wolf, S. Lamprier, O. Sigaud, and P.-Y. Oudeyer, “Grounding large language models in interactive environments with online reinforcement learning,” in International Conference on Machine Learning (ICML) . PMLR, 2023
2023
Cited alongside, same era.
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian et al. , “Do as I can, not as I say: Grounding language in robotic affordances,” in Conference on Robot Learning (CoRL) . PMLR, 2023
2023
Cited alongside, same era.
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, T. Jackson, N. Brown, L. Luu, S. Levine, K. Hausman, and B. Ichter, “Inner monologue: Embodied reasoning through planning with language models,” in Conference on Robot Learning (CoRL) . PMLR, 2023
2023
Cited alongside, same era.
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg, “ProgPrompt: Generating situated robot task plans using large language models,” in IEEE International Conference on Robotics and Automation (ICRA) , 2023
2023
Cited alongside, same era.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in IEEE International Conference on Robotics and Automation (ICRA) , 2023
2023
Cited alongside, same era.
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh, “Reward design with language models,” in International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
2024
Closest in time.
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar, “Eureka: Human-level reward design via coding large language models,” in International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
E. Triantafyllidis, F. Christianos, and Z. Li, “Intrinsic language-guided exploration for complex long-horizon robotic manipulation tasks,” in IEEE International Conference on Robotics and Automation (ICRA) , 2024
2024
Closest in time.
L. Chen, Y. Lei, S. Jin, Y. Zhang, and L. Zhang, “Rlingua: Improving reinforcement learning sample efficiency in robotic manipulations with large language models,” IEEE Robotics and Automation Letters , vol. 9, no. 7, pp. 6075–6082, 2024
2024
Closest in time.
B. van der Heijden, J. Luijkx, L. Ferranti, J. Kober, and R. Babuska, “Engine agnostic graph environments for robotics (EAGERx): A graph-based framework for sim2real robot learning,” IEEE Robotics and Automation Magazine , pp. 2–15, 2024
2024
Closest in time.
Delft AI Cluster (DAIC), “The Delft AI Cluster (DAIC), RRID:SCR_025091,” 2024. [Online]. Available: https://doc.daic.tudelft.nl/
2024
Closest in time.