Fetching the paper…
Reading the bibliography…
State-of-the-art (SOTA) reinforcement learning (RL) methods have enabled vision-language model (VLM) agents to learn from interaction with online environments without human supervision.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1992
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Variational inference for policy search in changing situations
G. Neumann · 2011
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Relative entropy regularized policy iteration
A. Abdolmaleki, J. T. Springenberg, J. Degrave, S. Bohez, Y. Tassa, D. Belov, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
N. Jiang and A. Agarwal · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dkebiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Earlier work this paper cites.
BabyAI: First steps towards grounded language learning with a human in the loop
M. Chevalier-Boisvert, D. Bahdanau, S. Lahlou, L. Willems, C. Saharia, T. H. Nguyen, and Y. Bengio · 2019
Earlier work this paper cites.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
S. Nair and C. Finn · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Sub-goal trees a framework for goal-based reinforcement learning
T. Jurgenson, O. Avner, E. Groshev, and A. Tamar · 2020
Earlier work this paper cites.
Divide-and-conquer monte carlo tree search for goal-directed planning
G. Parascandolo, L. Buesing, J. Merel, L. Hasenclever, J. Aslanides, J. B. Hamrick, N. Heess, A. Neitz, and T. Weber · 2020
Cited alongside, same era.
Goal-conditioned reinforcement learning with imagined subgoals
E. Chane-Sane, C. Schmid, and I. Laptev · 2021
Cited alongside, same era.
Androidenv: A reinforcement learning platform for android
D. Toyama, P. Hamel, A. Gergely, G. Comanici, A. Glaese, Z. Ahmed, T. Jackson, S. Mourad, and D. Precup · 2021
Cited alongside, same era.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Cited alongside, same era.
Constrained variational policy optimization for safe reinforcement learning
Z. Liu, Z. Cen, V. Isenbaev, W. Liu, S. Wu, B. Li, and D. Zhao · 2022
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
H. Bai, Y. Zhou, M. Cemri, J. Pan, A. Suhr, S. Levine, and A. Kumar · 2024
Later among the works it cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Later among the works it cites.
Cogagent: A visual language model for gui agents
W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, et al · 2024
Later among the works it cites.
Visualagentbench: Towards large multimodal models as visual foundation agents
X. Liu, T. Zhang, Y. Gu, I. L. Iong, Y. Xu, X. Song, S. Zhang, H. Lai, X. Liu, H. Zhao, et al · 2024
Later among the works it cites.
Vlp: Vision language planning for autonomous driving
C. Pan, B. Yaman, T. Nesti, A. Mallik, A. G. Allievi, S. Velipasalar, and L. Ren · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
S. Hong, X. Zheng, J. Chen, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, et al · 2023
Cited alongside, same era.
Gpt-4v(ision) technical work and authors
OpenAI · 2023
Cited alongside, same era.
Vision-language models are zero-shot reward models for reinforcement learning
J. Rocamonde, V. Montesinos, E. Nava, E. Perez, and D. Lindner · 2023
Cited alongside, same era.
Auto-gpt for online decision making: Benchmarks and additional opinions
H. Yang, S. Yue, and Y. He · 2023
Cited alongside, same era.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users
Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu · 2023
Cited alongside, same era.
Later among the works it cites.
Autonomous evaluation and refinement of digital agents
J. Pan, Y. Zhang, N. Tomlin, Y. Zhou, S. Levine, and A. Suhr · 2024
Later among the works it cites.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning
Z. Qi, X. Liu, I. L. Iong, H. Lai, X. Sun, X. Yang, J. Sun, Y. Yang, S. Yao, T. Zhang, et al · 2024
Later among the works it cites.
Androidinthewild: A large-scale dataset for android device control
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap · 2024
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
Distrl: An asynchronous distributed reinforcement learning framework for on-device control agents
T. Wang, Z. Wu, J. Liu, J. Hao, J. Wang, and K. Shao · 2024
Later among the works it cites.
Variational delayed policy optimization
Q. Wu, S. S. Zhan, Y. Wang, Y. Wang, C.-W. Lin, C. Lv, Q. Zhu, and C. Huang · 2024
Later among the works it cites.
Guiding long-horizon task and motion planning with vision language models
Z. Yang, C. Garrett, D. Fox, T. Lozano-Pérez, and L. P. Kaelbling · 2024
Later among the works it cites.
Epo: Hierarchical llm agents with environment preference optimization
Q. Zhao, H. Fu, C. Sun, and G. Konidaris · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Later among the works it cites.