Fetching the paper…
Reading the bibliography…
Large Language Model (LLM) agents, capable of performing a broad range of actions, such as invoking tools and controlling robots, show great potential in tackling real-world challenges.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and Liang, P · 2015
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Cote, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Do as i can and not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A · 2022
Earlier work this paper cites.
LangChain, October 2022
Chase, H · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities
Lee, M., Liang, P., and Yang, Q · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Sharegpt dataset
Anonymous · 2023
Earlier work this paper cites.
Access: Advancing innovation: Nsf’s advanced cyberinfrastructure coordination ecosystem: Services & support
Boerner, T. J., Deems, S., Furlani, T. R., Knuth, S. L., and Towns, J · 2023
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Bran, A. M., Cox, S., White, A. D., and Schwaller, P · 2023
Earlier work this paper cites.
epfllm megatron-llm, 2023
Cano, A. H., Pagliardini, M., Köpf, A., Matoba, K., Mohtashami, A., Wang, X., Fan, O. S., Marmet, A., Bayazit, D., Krawczuk, I., Chen, Z., Salvi, F., Bosselut, A., and Jaggi, M · 2023
Cited alongside, same era.
Gpts are gpts: An early look at the labor market impact potential of large language models
Eloundou, T., Manning, S., Mishkin, P., and Rock, D · 2023
Cited alongside, same era.
Reflective linguistic programming (rlp): A stepping stone in socially-aware agi (socialagi)
Fischer, K. A · 2023
Cited alongside, same era.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Hong, S., Zheng, X., Chen, J., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., et al · 2023
Taskweaver: A code-first agent framework
Qiao, B., Li, L., Zhang, X., He, S., Kang, Y., Zhang, C., Yang, F., Dong, H., Zhang, J., Wang, L., et al · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Fei-Fei, L · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Language models can solve computer tasks
Kim, G., Baldi, P., and McAleer, S · 2023
Cited alongside, same era.
Capybara dataset
LDJnr · 2023
Cited alongside, same era.
Api-bank: A benchmark for tool-augmented llms, 2023
Li, M., Song, F., Yu, B., Yu, H., Li, Z., Huang, F., and Li, Y · 2023
Cited alongside, same era.
Openorca: An open dataset of gpt augmented flan reasoning traces
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium” · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Cited alongside, same era.
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Code4Struct: Code generation for few-shot event structure prediction
Wang, X., Li, S., and Ji, H · 2023
Later among the works it cites.
On the tool manipulation capability of open-source large language models, 2023
Xu, Q., Hong, F., Li, B., Hu, C., Chen, Z., and Zhang, J · 2023
Later among the works it cites.
Craft: Customizing llms by creating and retrieving from specialized toolsets
Yuan, L., Chen, Y., Wang, X., Fung, Y. R., Peng, H., and Ji, H · 2023
Later among the works it cites.
Agenttuning: Enabling generalized agent abilities for llms, 2023
Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., and Tang, J · 2023
Later among the works it cites.
Prefer: Prompt ensemble learning via feedback-reflect-refine
Zhang, C., Liu, L., Wang, J., Wang, C., Sun, X., Wang, H., and Cai, M · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Zhu, X., Chen, Y., Tian, H., Tao, C., Su, W., Yang, C., Huang, G., Li, B., Lu, L., Wang, X., et al · 2023
Later among the works it cites.
Data interpreter: An llm agent for data science, 2024
Hong, S., Lin, Y., Liu, B., Liu, B., Wu, B., Li, D., Chen, J., Zhang, J., Wang, J., Zhang, L., Zhang, L., Yang, M., Zhuge, M., Guo, T., Zhou, T., Tao, W., Wang, W., Tang, X., Lu, X., Zheng, X., Liang, X., Fei, Y., Cheng, Y., Xu, Z., and Wu, C · 2024
Closest in time.
Prioritizing safeguarding over autonomy: Risks of llm agents for science
Tang, X., Jin, Q., Zhu, K., Yuan, T., Zhang, Y., Zhou, W., Qu, M., Zhao, Y., Tang, J., Zhang, Z., et al · 2024
Closest in time.
Tiobe index
TIOBE Index · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
Zheng, T., Zhang, G., Shen, T., Liu, X., Lin, B. Y., Fu, J., Chen, W., and Yue, X · 2024
Closest in time.