Fetching the paper…
Reading the bibliography…
AI Agents are changing the way work gets done, both in consumer and enterprise domains.
Generating project networks
Tate, A. (1977) · 1977
Earlier work this paper cites.
Shop: Simple hierarchical ordered planner
Nau, D., Cao, Y., Lotem, A., and Muñoz-Avila, H. (1991) · 1991
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Angelic semantics for high-level actions
Marthi, B., Russell, S., and Wolfe, J. (2007) · 2007
Earlier work this paper cites.
Artificial Intelligence: a modern approach
Russell, S. J. and Norvig, P. (2009) · 2009
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D. (2017) · 2017
Earlier work this paper cites.
Compositional generalization for natural language interfaces to web apis
Hosseini, S., Awadallah, A. H., and Su, Y. (2021) · 2021
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokhov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J. (2022) · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. (2022) · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Irpan, A., Herzog, A., Toshev, A. T., Zeng, A., Brohan, A., Ichter, B. A., David, B., Parada, C., Finn, C., Tan, C., Reyes, D., Kalashnikov, D., Jang, E. V., Xia, F., Rettinghouse, J. L., Hsu, J. C., Quiambao, J. L., Ibarz, J., Rao, K., Hausman, K., Gopalakrishnan, K., Lee, K.-H., Jeffrey, K. A., Luu, L., Yan, M., Ahn, M. S., Sievers, N., Joshi, N. J., Brown, N., Cortes, O. E. E., Xu, P., Sampedro, P. P., Sermanet, P., Ruano, R. J., Julian, R. C., Jesmonth, S. A., Levine, S., Xu, S., Xiao, T., Vanhoucke, V. O., Lu, Y., Chebotar, Y., and Kuang, Y. (2022) · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. (2022) · 2022
Earlier work this paper cites.
Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents
Wang, Z., Cai, S., Chen, G., Liu, A., Ma, X., and Liang, Y. (2022) · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K. (2022) · 2022
Cited alongside, same era.
Robotic offline rl from internet videos via value-function pre-training
Bhateja, C., Guo, D., Ghosh, D., Singh, A., Tomar, M., Vuong, Q., Chebotar, Y., Levine, S., and Kumar, A. (2023) · 2023
Cited alongside, same era.
Robocat: A self-improving foundation agent for robotic manipulation
Bousmalis, K., Vezzani, G., Rao, D., Devin, C., Lee, A. X., Bauza, M., Davchev, T., Zhou, Y., Gupta, A., Raju, A., Laurens, A., Fantacci, C., Dalibard, V., Zambelli, M., Martins, M., Pevceviciute, R., Blokzijl, M., Denil, M., Batchelor, N., Lampe, T., Parisotto, E., Żołna, K., Reed, S., Colmenarejo, S. G., Scholz, J., Abdolmaleki, A., Groth, O., Regli, J.-B., Sushkov, O., Rothörl, T., Chen, J. E., Aytar, Y., Barker, D., Ortiz, J., Riedmiller, M., Springenberg, J. T., Hadsell, R., Nori, F., and Heess, N. (2023) · 2023
Cited alongside, same era.
A survey of chain of thought reasoning: Advances, frontiers and future
Chu, Z., Chen, J., Chen, Q., Yu, W., He, T., Wang, H., Peng, W., Liu, M., Qin, B., and Liu, T. (2023) · 2023
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. (2024) · 2024
Closest in time.
Webvoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D. (2024) · 2024
Closest in time.
Llms can’t plan, but can help planning in llm-modulo frameworks
Kambhampati, S., Valmeekam, K., Guan, L., Verma, M., Stechly, K., Bhambri, S., Saldyt, L., and Murthy, A. (2024) · 2024
Closest in time.
Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., and Narayanan, A. (2024) · 2024
Closest in time.
Wilbur: Adaptive in-context learning for robust and accurate web agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T. (2023) · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D. (2023) · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5 technical report
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T. (2023) · 2023
Cited alongside, same era.
Evaluating cognitive maps and planning in large language models with cogeval
Momennejad, I., Hasanbeig, H., Vieira, F., Sharma, H., Ness, R. O., Jojic, N., Palangi, H., and Larson, J. (2023) · 2023
Cited alongside, same era.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E. (2023) · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023) · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., et al. (2023) · 2023
Cited alongside, same era.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Bai, H., Zhou, Y., Cemri, M., Pan, J., Suhr, A., Levine, S., and Kumar, A. (2024) · 2024
Cited alongside, same era.
Lutz, M., Bohra, A., Saroyan, M., Harutyunyan, A., and Campagna, G. (2024) · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A., Khanpour, H., Rosset, C., and Awadallah, A. (2024) · 2024
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., and Sun, M. (2024) · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S. (2024) · 2024
Closest in time.
Autodroid: Llm-powered task automation in android
Wen, H., Li, Y., Liu, G., Zhao, S., Yu, T., Li, T. J.-J., Jiang, S., Liu, Y., Zhang, Y., and Liu, Y. (2024) · 2024
Closest in time.
Automating the enterprise with foundation models
Wornow, M., Narayan, A., Opsahl-Ong, K., McIntyre, Q., Shah, N. H., and Re, C. (2024) · 2024
Closest in time.
Gpt-4v (ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y. (2024) · 2024
Closest in time.