Fetching the paper…
Reading the bibliography…
We propose the problem of conversational web navigation, where a digital agent controls a web browser and follows user instructions to solve real-world tasks in a multi-turn dialogue fashion.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
The distribution of the flora in the alpine zone. 1
Jaccard, P · 1912
Earlier work this paper cites.
A Survey of Web Information Extraction Systems
Chang, C., Kayed, M., Girgis, M. R., and Shaalan, K. F · 2006
Earlier work this paper cites.
ViDE: A Vision-based Approach for Deep Web Data Extraction
Liu, W., Meng, X., and Meng, W · 2009
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M. Z., Ranzato, M., Senior, A. W., Tucker, P. A., Yang, K., and Ng, A. Y · 2012
Earlier work this paper cites.
Usability Engineering
Carroll, J. M. and Rosson, M. B · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Popovic, M · 2015
Earlier work this paper cites.
A Diagram is Worth a Dozen Images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M. J., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
End-to-end Goal-driven Web Navigation
Nogueira, R. F. and Cho, K · 2016
Earlier work this paper cites.
A Survey on Dialogue Systems: Recent Advances and New Frontiers
Chen, H., Liu, X., Yin, D., and Tang, J · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Towards an Improved Vision-based Web Page Segmentation Algorithm
Cormer, M., Mann, R., Moffatt, K., and Cohen, R · 2017
Earlier work this paper cites.
World of Bits: An Open-domain Platform for Web-based Agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Reinforcement Learning on Web Interfaces using Workflow-guided Exploration
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Earlier work this paper cites.
Mapping natural language commands to web elements
Pasupat, P., Jiang, T., Liu, E. Z., Guu, K., and Liang, P · 2018
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y., Khot, T., Agrawal, A., Summers-Stay, D., Batra, D., and Parikh, D · 2019
Earlier work this paper cites.
Learning conversational web interfaces
Gur, I. and Yan, X · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Language Models are Few-shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Dense Passage Retrieval for Open-domain Question Answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P. S. H., Wu, L., Edunov, S., Chen, D., and Yih, W · 2020
Earlier work this paper cites.
Web Page Segmentation Revisited: Evaluation Framework and Dataset
Kiesel, J., Kneist, F., Meyer, L., Komlossy, K., Stein, B., and Potthast, M · 2020
Earlier work this paper cites.
Mapping Natural Language Instructions to Mobile UI Action Sequences
Li, Y., He, J., Zhou, X., Zhang, Y., and Baldridge, J · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Towards Scalable Multi-domain Conversational Agents: The Schema-guided Dialogue Dataset
Rastogi, A., Zang, X., Sunkara, S., Gupta, R., and Khaitan, P · 2020
Earlier work this paper cites.
MiniLM: Deep Self-attention Distillation for Task-agnostic Compression of Pre-trained Transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M · 2020
Earlier work this paper cites.
DIALOGPT : Large-scale Generative Pre-training for Conversational Response Generation
Zhang, Y., Sun, S., Galley, M., Chen, Y., Brockett, C., Gao, X., Gao, J., Liu, J., and Dolan, B · 2020
Earlier work this paper cites.
VINS: Visual Search for Mobile User Interface Design
Bunian, S., Li, K., Jemmali, C., Harteveld, C., Fu, Y., and El-Nasr, M. S · 2021
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Environment Generation for Zero-shot Compositional Reinforcement Learning
Gur, I., Jaques, N., Miao, Y., Choi, J., Tiwari, M., Lee, H., and Faust, A · 2021
Cited alongside, same era.
Deberta: decoding-enhanced Bert with Disentangled Attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Cited alongside, same era.
Multimodal Web Navigation with Instruction-finetuned Foundation Models
Furuta, H., Nachum, O., Lee, K., Matsuo, Y., Gu, S. S., and Gur, I · 2023
Later among the works it cites.
The bfloat16 numerical format — Cloud TPU, December 2023
Google · 2023
Later among the works it cites.
Are LLMs All You Need for Task-oriented Dialogue?
Hudeček, V. and Dušek, O · 2023
Later among the works it cites.
HyperWrite AI Personal Assistant – ”The first publicly available AI agent that can operate a browser like a human.”
HyperWrite · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2021
Cited alongside, same era.
Introducing Triton: Open-source GPU programming for neural networks, July 2021
OpenAI · 2021
Cited alongside, same era.
Screen Parsing: Towards Reverse Engineering of UI Models from Screenshots
Wu, J., Zhang, X., Nichols, J., and Bigham, J. P · 2021
Cited alongside, same era.
Grounding Open-domain Instructions to Automate Web Support Tasks
Xu, N., Masling, S., Du, M., Campagna, G., Heck, L., Landay, J. A., and Lam, M · 2021
Cited alongside, same era.
Simplified DOM Trees for Transferable Attribute Extraction from the Web
Zhou, Y., Sheng, Y., Vo, N., Edmonds, N., and Tata, S · 2021
Cited alongside, same era.
TopiOCQA: Open-domain Conversational Question Answering with Topic Switching
Adlakha, V., Dhuliawala, S., Suleman, K., de Vries, H., and Reddy, S · 2022
Cited alongside, same era.
HTLM: Hyper-text Pre-training and Prompting of Language Models
Aghajanyan, A., Okhonko, D., Lewis, M., Joshi, M., Xu, H., Ghosh, G., and Zettlemoyer, L · 2022
Cited alongside, same era.
Later among the works it cites.
Language Models can Solve Computer Tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Later among the works it cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Later among the works it cites.
OBELICS: An Open Web-scale Filtered Dataset of Interleaved Image-text Documents, 2023
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., Cord, M., and Sanh, V · 2023
Later among the works it cites.
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M., and Toutanova, K · 2023
Later among the works it cites.
MTEB: Massive Text Embedding Benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2023
Later among the works it cites.
Multi-on – ”The world’s first Personal AI Agent & Life Copilot”
Multi-On · 2023
Later among the works it cites.
ChatGPT Plugins
OpenAI · 2023
Later among the works it cites.
GPT-3.5 Turbo fine-tuning and API updates, August 2023
Peng, A., Wu, M., Kilpatrick, L., and Heidel, S · 2023
Later among the works it cites.
Bard can now connect to your google apps and services, Sep 2023
Pinsky, Y · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Android in the Wild: A Large-scale Dataset for Android Device Control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T. P · 2023
Later among the works it cites.
From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
Shaw, P., Joshi, M., Cohan, J., Berant, J., Pasupat, P., Hu, H., Khandelwal, U., Lee, K., and Toutanova, K · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models, october 2023
Together · 2023
Later among the works it cites.
WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics
Wu, J., Wang, S., Shen, S., Peng, Y., Nichols, J., and Bigham, J. P · 2023
Later among the works it cites.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2023
Later among the works it cites.
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., and Neubig, G · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-language Understanding with Advanced Large Language Models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Workarena: How capable are web agents at solving common knowledge work tasks?, 2024
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Boisvert, L., Thakkar, M., Cappart, Q., Vazquez, D., Chapados, N., and Lacoste, A · 2024
Closest in time.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A. V., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2024
Closest in time.
WebVoyager: Building an End-to-end Web Agent with Large Multimodal Models, 2024
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Closest in time.
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks, 2024
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Closest in time.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T · 2024
Closest in time.
Tur[k]ingbench: A challenge benchmark for web agents, 2024
Xu, K., Kordi, Y., Sanders, K., Wang, Y., Byerly, A., Zhang, J., Durme, B. V., and Khashabi, D · 2024
Closest in time.
GPT-4V(ision) is a Generalist Web Agent, if Grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Closest in time.
Recent advances and challenges in task-oriented dialog systems
Zhang, Z., Takanobu, R., Zhu, Q., Huang, M., and Zhu, X · 2027
Closest in time.