Fetching the paper…
Reading the bibliography…
We interact with computers on an everyday basis, be it in everyday life or work, and many aspects of work can be done entirely with access to a computer and the Internet.
Development of occupational interest profiles for o* net
Rounds, J., Smith, T., Hubert, L., Lewis, P., and Rivkin, D · 1999
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku, 2023
Anthropic · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Earlier work this paper cites.
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models, 2023
Eloundou, T., Manning, S., Mishkin, P., and Rock, D · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior, 2023
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Earlier work this paper cites.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Earlier work this paper cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2023
Earlier work this paper cites.
Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity | Lex Fridman Podcast #452, November 2024
Amodei, D. and Fridman, L · 2024
Earlier work this paper cites.
Windows agent arena: Evaluating multi-modal os agents at scale
Bonatti, R., Zhao, D., Bonacci, F., Dupont, D., Abdali, S., Li, Y., Lu, Y., Wagle, J., Koishida, K., Bucker, A., et al · 2024
Earlier work this paper cites.
The browsergym ecosystem for web agent research
Chezelles, D., Le Sellier, T., Gasse, M., Lacoste, A., Drouin, A., Caccia, M., Boisvert, L., Thakkar, M., Marty, T., Assouel, R., et al · 2024
Cited alongside, same era.
ARC Prize 2024: Technical Report, December 2024
Chollet, F., Knoop, M., Kamradt, G., and Landers, B · 2024
Cited alongside, same era.
Workarena: How capable are web agents at solving common knowledge work tasks?, 2024
Drouin, A., Gasse, M., Caccia, M., Laradji, I. H., Verme, M. D., Marty, T., Boisvert, L., Thakkar, M., Cappart, Q., Vazquez, D., Chapados, N., and Lacoste, A · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
Weblinx: Real-world website navigation with multi-turn dialogue, 2024
Lù, X. H., Kasner, Z., and Reddy, S · 2024
Closest in time.
The 29.1 release of the O*NET database, November 2024
O*NET · 2024
Closest in time.
Introducing gpt-4o and more tools to chatgpt free users, 2024
OpenAI · 2024
Closest in time.
Beyond browsing: Api-based web agents
Song, Y., Xu, F., Zhou, S., and Neubig, G · 2024
Closest in time.
AppWorld: A controllable world of apps and people for benchmarking interactive coding agents
Trivedi, H., Khot, T., Hartmann, M., Manku, R., Dong, V., Li, E., Gupta, S., Sabharwal, A., and Balasubramanian, N · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Cited alongside, same era.
Huang, K.-H., Prabhakar, A., Dhawan, S., Mao, Y., Wang, H., Savarese, S., Xiong, C., Laban, P., and Wu, C.-S · 2024
Cited alongside, same era.
The amazon nova family of models: Technical report and model card
Intelligence, A. A. G · 2024
Cited alongside, same era.
Videowebarena: Evaluating long context multimodal agents with video understanding web tasks, 2024
Jang, L., Li, Y., Ding, C., Lin, J., Liang, P. P., Zhao, D., Bonatti, R., and Koishida, K · 2024
Cited alongside, same era.
SWE-bench: Can Language Models Resolve Real-world Github Issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Cited alongside, same era.
Llms can’t plan, but can help planning in llm-modulo frameworks
Kambhampati, S., Valmeekam, K., Guan, L., Verma, M., Stechly, K., Bhambri, S., Saldyt, L., and Murthy, A · 2024
Cited alongside, same era.
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Cited alongside, same era.
Devbench: A comprehensive benchmark for software development
Li, B., Wu, W., Tang, Z., Shi, L., Yang, J., Li, J., Yao, S., Qian, C., Hui, B., Zhang, Q., et al · 2024
Cited alongside, same era.
Wittenstein, J · 2024
Closest in time.
OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains
Yao, S., Shinn, N., Razavi, P., and Narasimhan, K · 2024
Closest in time.
Assistantbench: Can web agents solve realistic and time-consuming tasks?, 2024
Yoran, O., Amouyal, S. J., Malaviya, C., Bogin, B., Press, O., and Berant, J · 2024
Closest in time.
SOTOPIA: Interactive evaluation for social intelligence in language agents
Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y., Fried, D., Neubig, G., and Sap, M · 2024
Closest in time.
Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation, 2025
Hu, M., Zhou, Y., Fan, W., Nie, Y., Xia, B., Sun, T., Ye, Z., Jin, Z., Li, Y., Zhang, Z., Wang, Y., Ye, Q., Luo, P., and Li, G · 2025
Closest in time.