Fetching the paper…
Reading the bibliography…
Goal-conditioned and Multi-Task Reinforcement Learning (GCRL and MTRL) address numerous problems related to robot learning, including locomotion, navigation, and manipulation scenarios.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. M. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 1910
Earlier work this paper cites.
Robot shaping: Developing autonomous agents through learning
M. Dorigo and M. Colombetti · 1994
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
J. Randløv and P. Alstrøm · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2005
Earlier work this paper cites.
Language-conditioned goal generation: a new approach to language grounding for rl
C. Colas, A. Akakzia, P.-Y. Oudeyer, M. Chetouani, and O. Sigaud · 2006
Earlier work this paper cites.
Language-conditioned goal generation: a new approach to language grounding for RL
C. Colas, A. Akakzia, P. Oudeyer, M. Chetouani, and O. Sigaud · 2006
Earlier work this paper cites.
A survey of deep network solutions for learning control in robotics: From reinforcement to imitation
L. Tai, J. Zhang, M. Liu, J. Boedecker, and W. Burgard · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. M. O. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Feature-based transfer learning for robotic push manipulation
J. Stüber, M. Kopicki, and C. Zito · 2018
Cited alongside, same era.
Hg-dagger: Interactive imitation learning with human experts
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, V. Kumar, and W. Zaremba · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. Nair, V. H. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Multi-modal transfer learning for grasping transparent and specular objects
T. Weng, A. Pallankize, Y. Tang, O. Kroemer, and D. Held · 2020
Cited alongside, same era.
Vima: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2022
Later among the works it cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osinski, B. Ichter, and S. Levine · 2022
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Later among the works it cites.
Should i run offline reinforcement learning or behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interactive reinforcement learning with inaccurate feedback
T. A. K. Faulkner, E. S. Short, and A. L. Thomaz · 2020
Cited alongside, same era.
Transfer learning for accurate modeling and control of soft actuators
M. Wiese, G. Runge-Borchert, B.-H. Cao, and A. Raatz · 2021
Cited alongside, same era.
Correct me if i am wrong: Interactive learning for robotic manipulation
E. Chisari, T. Welschehold, J. Boedecker, W. Burgard, and A. Valada · 2021
Cited alongside, same era.
Asymmetric self-play for automatic goal discovery in robotic manipulation
O. OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. de Oliveira Pinto, A. Paino, H. Noh, L. Weng, Q. Yuan, C. Chu, and W. Zaremba · 2021
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter · 2022
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020a
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu
Cited in the paper.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. A. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. Bang, A. Madotto, and P. Fung · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Reward design with language models
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh · 2023
Closest in time.
Starcoder: may the source be with you!, 2023
R. Li, L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, C. Mou, M. Marone, C. Akiki, J. Li, J. Chim, Q. Liu, E. Zheltonozhskii, T. Y. Zhuo, T. Wang, O. Dehaene, M. Davaadorj, J. Lamy-Poirier, J. Monteiro, O. Shliazhko, N. Gontier, N. Meade, A. Zebaze, M.-H. Yee, L. K. Umapathi, J. Zhu, B. Lipkin, M. Oblokulov, Z. Wang, R. Murthy, J. Stillerman, S. S. Patel, D. Abulkhanov, M. Zocca, M. Dey, Z. Zhang, N. Fahmy, U. Bhattacharyya, W. Yu, S. Singh, S. Luccioni, P. Villegas, M. Kunakov, F. Zhdanov, M. Romero, T. Lee, N. Timor, J. Ding, C. Schlesinger, H. Schoelkopf, J. Ebert, T. Dao, M. Mishra, A. Gu, J. Robinson, C. J. Anderson, B. Dolan-Gavitt, D. Contractor, S. Reddy, D. Fried, D. Bahdanau, Y. Jernite, C. M. Ferrandis, S. Hughes, T. Wolf, A. Guha, L. von Werra, and H. de Vries · 2023
Closest in time.