Fetching the paper…
Reading the bibliography…
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion.
The theory of dynamic programming
Bellman, R · 1954
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Multiple model-based reinforcement learning
Doya, K., Samejima, K., Katagiri, K.-i., and Kawato, M · 2002
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Earlier work this paper cites.
When waiting is not an option : Learning options with a deliberation cost
Harb, J., Bacon, P.-L., Klissarov, M., and Precup, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Simplifying reward design through divide-and-conquer, 2018
Ratner, E., Hadfield-Menell, D., and Dragan, A. D · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Dealing with sparse rewards in reinforcement learning
Hare, J · 2019
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Earlier work this paper cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Kuznetsov, A., Shvechikov, P., Grishin, A., and Vetrov, D · 2020
Earlier work this paper cites.
Pitfalls of learning a reward function online
Armstrong, S., Leike, J., Orseau, L., and Legg, S · 2021
Earlier work this paper cites.
Explicable reward design for reinforcement learning agents
Devidze, R., Radanovic, G., Kamalaruban, P., and Singla, A · 2021
Earlier work this paper cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., and Hester, T · 2021
Cited alongside, same era.
panda-gym: Open-Source Goal-Conditioned Environments for Robotic Learning
Gallouédec, Q., Cazin, N., Dellandréa, E., and Chen, L · 2021
Cited alongside, same era.
Flexible option learning
Klissarov, M. and Precup, D · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Learning about progress from experts
Bruce, J., Anand, A., Mazoure, B., and Fergus, R · 2023
Later among the works it cites.
Chevalier-Boisvert, M., Dai, B., Towers, M., de Lazcano, R., Willems, L., Lahlou, S., Pal, S., Castro, P. S., and Terry, J · 2023
Later among the works it cites.
Clip4mc: An rl-friendly vision-language model for minecraft
Ding, Z., Luo, H., Li, K., Yue, J., Huang, T., and Lu, Z · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback, 2023
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M · 2023
Later among the works it cites.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R. C., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D. M., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., and Yan, M · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning, 2022
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Can foundation models perform zero-shot task specification for robot manipulation?
Cui, Y., Niekum, S., Gupta, A., Kumar, V., and Rajeswaran, A · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L. J., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Brown, N., Jackson, T., Luu, L., Levine, S., Hausman, K., and Ichter, B · 2022
Cited alongside, same era.
Minerl diamond 2021 competition: Overview, results, and lessons learned
Kanervisto, A., Milani, S., Ramanauskas, K., Topin, N., Lin, Z., Li, J., yong Shi, J., Ye, D., Fu, Q., Yang, W., Hong, W., Huang, Z.-H., Chen, H., Zeng, G., Lin, Y., Micheli, V., Alonso, E., Fleuret, F., Nikulin, A., Belousov, Y., Svidchenko, O., and Shpilman, A · 2022
Cited alongside, same era.
Zero-shot reward specification via grounded natural language
Mahmoudieh, P., Pathak, D., and Darrell, T · 2022
Cited alongside, same era.
Later among the works it cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Accelerating exploration and representation learning with offline pre-training
Mazoure, B., Bruce, J., Precup, D., Fergus, R., and Anand, A · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI, :, Achiam, J., and et al · 2023
Later among the works it cites.
Vision-language models are zero-shot reward models for reinforcement learning, 2023
Rocamonde, J., Montesinos, V., Nava, E., Perez, E., and Lindner, D · 2023
Later among the works it cites.
Roboclip: One demonstration is enough to learn robot policies
Sontakke, S. A., Zhang, J., Arnold, S. M. R., Pertsch, K., Biyik, E., Sadigh, D., Finn, C., and Itti, L · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models, 2023
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Gonzalez Arenas, M., Lewis Chiang, H.-T., Erez, T., Hasenclever, L., Humplik, J., Ichter, B., Xiao, T., Xu, P., Zeng, A., Zhang, T., Heess, N., Sadigh, D., Tan, J., Tassa, Y., and Xia, F · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.