Fetching the paper…
Reading the bibliography…
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-tasks within web-based environments.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Plow: A collaborative task learning agent
James Allen, Nathanael Chambers, George Ferguson, Lucian Galescu, Hyuckchul Jung, Mary Swift, and William Taysom. 2007 · 2007
Earlier work this paper cites.
Driving semantic parsing from the world’s response
James Clarke, Dan Goldwasser, Ming-Wei Chang, and Dan Roth. 2010 · 2010
Earlier work this paper cites.
Probabilistic frame-semantic parsing
Dipanjan Das, Nathan Schneider, Desai Chen, and Noah A Smith. 2010 · 2010
Earlier work this paper cites.
The Turking Test: Can Language Models Understand Instructions?
Avia Efrat and Omer Levy. 2020 · 2010
Earlier work this paper cites.
Joint learning of words and meaning representations for open-text semantic parsing
Antoine Bordes, Xavier Glorot, Jason Weston, and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Learning to navigate the web
Izzeddin Gur, Ulrich Rueckert, Aleksandra Faust, and Dilek Hakkani-Tur. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
About three-in-ten us adults say they are ‘almost constantly’online
Andrew Perrin and Madhu Kumar. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. 2020 · 2020
Earlier work this paper cites.
Mapping natural language instructions to mobile ui action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020 · 2020
Earlier work this paper cites.
HTLM: Hyper-Text Pre-Training and Prompting of Language Models
A. Aghajanyan et al. 2021 · 2021
Earlier work this paper cites.
UIBert: Learning generic multimodal representations for ui understanding
Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, et al. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, and Michael Petrov et al. 2021 · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021 · 2021
Earlier work this paper cites.
VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling
Yang Li, Gang Li, Xin Zhou, Mostafa Dehghani, and Alexey Gritsenko. 2021 · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Earlier work this paper cites.
AndroidEnv: A reinforcement learning platform for android
Daniel Toyama, Philippe Hamel, Anita Gergely, Gheorghe Comanici, Amelia Glaese, Zafarali Ahmed, Tyler Jackson, Shibl Mourad, and Doina Precup. 2021 · 2021
Earlier work this paper cites.
Grounding open-domain instructions to automate web support tasks
Nancy Xu, Sam Masling, Michael Du, Giovanni Campagna, Larry Heck, James Landay, and Monica Lam. 2021 · 2021
Earlier work this paper cites.
CM3: A causal masked multimodal model of the internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, et al. 2022 · 2022
Earlier work this paper cites.
A dataset for interactive vision-language navigation with unknown command feasibility
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A Plummer. 2022 · 2022
Cited alongside, same era.
XDoc: Unified Pre-training for Cross-Format Document Understanding
Jingye Chen, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Understanding html with large language models
Izzeddin Gur, Ofir Nachum, Yingjie Miao, Mustafa Safdari, Austin Huang, Aakanksha Chowdhery, Sharan Narang, Noah Fiedel, and Aleksandra Faust. 2022 · 2022
Cited alongside, same era.
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Tool learning with foundation models
Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Yufei Huang, Chaojun Xiao, Chi Han, et al. 2023 · 2023
Later among the works it cites.
Language modelling with pixels
Phillip Rust, Jonas F Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott. 2023 · 2023
Later among the works it cites.
ToolFormer: language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Hierarchical prompting assists large language model on web navigation
Abishek Sridhar, Robert Lo, Frank F Xu, Hao Zhu, and Shuyan Zhou. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter C Humphreys, David Raposo, Tobias Pohlen, Gregory Thornton, Rachita Chhaparia, Alistair Muldal, Josh Abramson, Petko Georgiev, Adam Santoro, and Timothy Lillicrap. 2022 · 2022
Cited alongside, same era.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu, Julian Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2022 · 2022
Cited alongside, same era.
Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus
Gang Li and Yang Li. 2022 · 2022
Cited alongside, same era.
MUG: Interactive Multimodal Grounding on User Interfaces
Tao Li, Gang Li, Jingjie Zheng, Purple Wang, and Yang Li. 2022 · 2022
Cited alongside, same era.
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu. 2022 · 2022
Cited alongside, same era.
Knowing where and what: Unified word block pretraining for document understanding
Song Tao, Zijian Wang, Tiantian Fan, Canjie Luo, and Can Huang. 2022 · 2022
Cited alongside, same era.
WebShop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022 · 2022
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
Heyi Tao, Sethuraman TV, Michal Shlapentokh-Rothman, Derek Hoiem, and Heng Ji. 2023 · 2023
Later among the works it cites.
WebArena: a realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al. 2023 · 2023
Later among the works it cites.
InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024 · 2024
Closest in time.
Mind2Web: towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024 · 2024
Closest in time.
Dual-view visual contextualization for web navigation
Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao. 2024 · 2024
Closest in time.
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. 2024 · 2024
Closest in time.
Visualagentbench: Towards large multimodal models as visual foundation agents
Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, et al. 2024 · 2024
Closest in time.
Jiarui Lu, Thomas Holleis, Yizhe Zhang, Bernhard Aumayer, Feng Nan, Felix Bai, Shuang Ma, Shen Ma, Mengyu Li, Guoli Yin, et al. 2024 · 2024
Closest in time.
Weblinx: Real-world website navigation with multi-turn dialogue
Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024 · 2024
Closest in time.
GEAR: augmenting language models with generalizable and efficient tool resolution
Yining Lu, Haoping Yu, and Daniel Khashabi. 2024 · 2024
Closest in time.
WebCanvas: benchmarking web agents in online environments
Yichen Pan, Dehan Kong, Sida Zhou, Cheng Cui, Yifei Leng, Bing Jiang, Hangyu Liu, Yanyi Shang, Shuyan Zhou, Tongshuang Wu, et al. 2024 · 2024
Closest in time.
AppWorld: a controllable world of apps and people for benchmarking interactive coding agents
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. 2024 · 2024
Closest in time.
OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. 2024 · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, et al. 2024 · 2024
Closest in time.