Fetching the paper…
Reading the bibliography…
LLMs can now act as autonomous agents that interact with digital environments and complete specific objectives (e.g., arranging an online meeting).
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
Visualizing data using t-sne
L. van der Maaten and G. E. Hinton · 2008
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
S. Branavan, H. Chen, L. Zettlemoyer, and R. Barzilay · 2009
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Reinforcement cutting-agent learning for video object segmentation
J. Han, L. Yang, D. Zhang, X. Chang, and X. Liang · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang · 2018
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Earlier work this paper cites.
Reasoning about goals, steps, and temporal ordering with wikihow
L. Zhang, Q. Lyu, and C. Callison-Burch · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
M. Shridhar, X. Yuan, M. Côté, Y. Bisk, A. Trischler, and M. J. Hausknecht · 2021
Earlier work this paper cites.
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Earlier work this paper cites.
A data-driven approach for learning to control computers
P. C. Humphreys, D. Raposo, T. Pohlen, G. Thornton, R. Chhaparia, A. Muldal, J. Abramson, P. Georgiev, A. Santoro, and T. Lillicrap · 2022
Earlier work this paper cites.
Language models of code are few-shot commonsense learners
A. Madaan, S. Zhou, U. Alon, Y. Yang, and G. Neubig · 2022
Earlier work this paper cites.
Clueweb22: 10 billion web documents with visual and semantic information
A. Overwijk, C. Xiong, X. Liu, C. VandenBerg, and J. Callan · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Guess the instruction! flipped learning makes language models stronger zero-shot learners
S. Ye, D. Kim, J. Jang, J. Shin, and M. Seo · 2022
Cited alongside, same era.
Hierarchical control of situated agents through natural language
S. Zhou, P. Yin, and G. Neubig · 2022
Cited alongside, same era.
Show me more details: Discovering hierarchies of procedures from semi-structured web data
S. Zhou, L. Zhang, Y. Yang, Q. Lyu, P. Yin, C. Callison-Burch, and G. Neubig · 2022
Cited alongside, same era.
Fireact: Toward language agent fine-tuning
B. Chen, C. Shu, E. Shareghi, N. Collier, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web, 2023
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users, 2023
C. Zhang, Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu · 2023
Later among the works it cites.
Agent-flan: Designing data and methods of effective agent tuning for large language models
Z. Chen, K. Liu, Q. Wang, W. Zhang, J. Liu, D. Lin, K. Chen, and F. Zhao · 2024
Closest in time.
Workarena: How capable are web agents at solving common knowledge work tasks?, 2024
A. Drouin, M. Gasse, M. Caccia, I. H. Laradji, M. D. Verme, T. Marty, L. Boisvert, M. Thakkar, Q. Cappart, D. Vazquez, N. Chapados, and A. Lacoste · 2024
Closest in time.
Webvoyager: Building an end-to-end web agent with large multimodal models
H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu · 2024
Closest in time.
Autowebglm: Bootstrap and reinforce a large language model-based web navigating agent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Furuta, O. Nachum, K.-H. Lee, Y. Matsuo, S. S. Gu, and I. Gur · 2023
Cited alongside, same era.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
I. Gur, H. Furuta, A. Huang, M. Safdari, Y. Matsuo, D. Eck, and A. Faust · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu · 2023
Cited alongside, same era.
Language models can solve computer tasks
G. Kim, P. Baldi, and S. McAleer · 2023
Cited alongside, same era.
Self-alignment with instruction backtranslation
X. Li, P. Yu, C. Zhou, T. Schick, L. Zettlemoyer, O. Levy, J. Weston, and M. Lewis · 2023
Cited alongside, same era.
Hierarchical prompting assists large language model on web navigation
C.-f. Lo, A. Sridhar, H. Zhu, F. F. Xu, and S. Zhou · 2023
Cited alongside, same era.
H. Lai, X. Liu, I. L. Iong, S. Yao, Y. Chen, P. Shen, H. Yu, H. Zhang, X. Zhang, Y. Dong, et al · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2024
Closest in time.
Bagel: Bootstrapping agents by guiding exploration with language
S. Murty, C. Manning, P. Shaw, M. Joshi, and K. Lee · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents
J. Pan, Y. Zhang, N. Tomlin, Y. Zhou, S. Levine, and A. Suhr · 2024
Closest in time.
Code prompting elicits conditional reasoning abilities in text+ code llms
H. Puerto, M. Tutek, S. Aditya, X. Zhu, and I. Gurevych · 2024
Closest in time.
From pixels to ui actions: Learning to follow instructions via graphical user interfaces
P. Shaw, M. Joshi, J. Cohan, J. Berant, P. Pasupat, H. Hu, U. Khandelwal, K. Lee, and K. N. Toutanova · 2024
Closest in time.
Design2code: How far are we from automating front-end engineering?
C. Si, Y. Zhang, Z. Yang, R. Liu, and D. Yang · 2024
Closest in time.
Trial and error: Exploration-based trajectory optimization for llm agents
Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin · 2024
Closest in time.
Executable code actions elicit better llm agents, 2024
X. Wang, Y. Chen, L. Yuan, Y. Zhang, Y. Li, H. Peng, and H. Ji · 2024
Closest in time.
Os-copilot: Towards generalist computer agents with self-improvement
Z. Wu, C. Han, Z. Ding, Z. Weng, Z. Liu, S. Yao, T. Yu, and L. Kong · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Closest in time.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig · 2024
Closest in time.