Fetching the paper…
Reading the bibliography…
Robots equipped with reinforcement learning (RL) have the potential to learn a wide range of skills solely from a reward signal.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
J. Fu, K. Luo, and S. Levine, “Learning robust rewards with adverserial inverse reinforcement learnin,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Xie, A. Singh, S. Levine, and C. Finn, “Few-shot goal inference for visuomotor learning and planning,” in Conference on Robot Learning , 2018, pp. 40–52
2018
Earlier work this paper cites.
A. Singh, L. Yang, K. Hartikainen, C. Finn, and S. Levine, “End-to-end robotic reinforcement learning without reward engineering,” in Robotics: Science and Systems , 2019
2019
Earlier work this paper cites.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 1179–1191, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Gupta, J. Yu, T. Z. Zhao, V. Kumar, A. Rovinsky, K. Xu, T. Devlin, and S. Levine, “Reset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention,” in 2021 IEEE International Conference on Robotics and Automation , 2021, pp. 6664–6671
2021
Earlier work this paper cites.
A. Kumar, A. Singh, F. Ebert, M. Nakamoto, Y. Yang, C. Finn, and S. Levine, “Pre-training for robots: Offline rl enables learning new tasks in a handful of trials,” in Robotics: Science and Systems , 2022
2022
Earlier work this paper cites.
M. S. Mark, A. Ghadirzadeh, X. Chen, and C. Finn, “Fine-tuning offline policies with optimistic action selection,” in Deep Reinforcement Learning Workshop NeurIPS 2022 , 2022
2022
Earlier work this paper cites.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman et al. , “Do as i can, not as i say: Grounding language in robotic affordances,” in Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter, “Inner monologue: Embodied reasoning through planning with language models,” in Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
D. Shah, B. Osinski, B. Ichter, and S. Levine, “LM-nav: Robotic navigation with large pre-trained models of language, vision, and action,” in 6th Annual Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
P. Mahmoudieh, D. Pathak, and T. Darrell, “Zero-shot reward specification via grounded natural language,” in Proceedings of the 39th International Conference on Machine Learning , vol. 162. PMLR, 2022, pp. 14 743–14 752
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. S. Raman, V. Cohen, E. Rosen, I. Idrees, D. Paulius, and S. Tellex, “Planning with large language models via corrective re-prompting,” in Foundation Models for Decision Making Workshop at NeurIPS 2022 , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta, “R3m: A universal visual representation for robot manipulation,” in Conference on Robot Learning , 2022
2022
Cited alongside, same era.
A. Sharma, K. Xu, N. Sardana, A. Gupta, K. Hausman, S. Levine, and C. Finn, “Autonomous reinforcement learning: Formalism and benchmarking,” in International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
C. Sun, J. Orbik, C. M. Devin, B. H. Yang, A. Gupta, G. Berseth, and S. Levine, “Fully autonomous real-world reinforcement learning with applications to mobile manipulation,” in Proceedings of the 5th Conference on Robot Learning , ser. Proceedings of Machine Learning Research, vol. 164. PMLR, 2022, pp. 308–319
2022
Cited alongside, same era.
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine, “Bridge data: Boosting generalization of robotic skills with cross-domain datasets,” in Robotics: Science and Systems , 2022
2022
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Balsells, M. Torne, Z. Wang, S. Desai, P. Agrawal, and A. Gupta, “Autonomous robotic reinforcement learning with asynchronous human feedback,” in Conference on Robot Learning , 2023
2023
Later among the works it cites.
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine, “Bridgedata v2: A dataset for robot learning at scale,” in Conference on Robot Learning , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler, “Open-vocabulary queryable scene representations for real world planning,” in IEEE International Conference on Robotics and Automation . IEEE, 2023, pp. 11 509–11 522
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan, “Vima: General robot manipulation with multimodal prompts,” in Fortieth International Conference on Machine Learning , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K.-H. Lee, M. G. Arenas, H.-T. L. Chiang, T. Erez, L. Hasenclever, J. Humplik, B. Ichter, T. Xiao, P. Xu, A. Zeng, T. Zhang, N. Heess, D. Sadigh, J. Tan, Y. Tassa, and F. Xia, “Language to rewards for robotic skill synthesis,” in Conference on Robot Learning , 2023
2023
Cited alongside, same era.
A. Sharma, A. M. Ahmed, R. Ahmad, and C. Finn, “Self-improving robots: End-to-end autonomous visuomotor reinforcement learning,” in 7th Annual Conference on Robot Learning , 2023
2023
Later among the works it cites.
M. Nakamoto, Y. Zhai, A. Singh, M. S. Mark, Y. Ma, C. Finn, A. Kumar, and S. Levine, “Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,” in Advances in Neural Information Processing Systems , 2023
2023
Later among the works it cites.
J. Yang, M. S. Mark, B. Vu, A. Sharma, J. Bohg, and C. Finn, “Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,” in 2024 IEEE International Conference on Robotics and Automation . IEEE, 2024
2024
Closest in time.
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia, “Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,” in Conference on Computer Vision and Pattern Recognition , 2024
2024
Closest in time.
2024
Closest in time.
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar, “Eureka: Human-level reward design via coding large language models,” in International Conference on Learning Representations , 2024
2024
Closest in time.
OpenAI, “Gpt-4o system card,” in arXiv preprint arXiv:2410.21276 , 2024
2024
Closest in time.
2024
Closest in time.
Y. Wang, Z. Sun, J. Zhang, Z. Xian, E. Biyik, D. Held, and Z. Erickson, “Rl-vlm-f: Reinforcement learning from vision language foundation model feedback,” in Proceedings of the 41th International Conference on Machine Learning , 2024
2024
Closest in time.
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei, “Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,” in Conference on Robot Learning , 2024
2024
Closest in time.
2024
Closest in time.
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” in International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “Cotracker: It is better to track together,” in European Conference on Computer Vision , 2024
2024
Closest in time.
K. Fang, F. Liu, P. Abbeel, and S. Levine, “Moka: Open-world robotic manipulation through mark-based visual prompting,” in Robotics: Science and Systems , 2024
2024
Closest in time.