Fetching the paper…
Reading the bibliography…
Agents based on large language models (LLMs) have demonstrated effectiveness in solving a wide range of tasks by integrating LLMs with key modules such as planning, memory, and tool usage.
Situations, actions, and causal laws
McCarthy, J., et al · 1963
Earlier work this paper cites.
Shakey the robot
Nilsson, N. J., et al · 1984
Earlier work this paper cites.
Prodigy: An integrated architecture for planning and learning
Carbonell, J., Etzioni, O., Gil, Y., Joseph, R., Knoblock, C., Minton, S., and Veloso, M · 1991
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Automated Planning: theory and practice
Ghallab, M., Nau, D., and Traverso, P · 2004
Earlier work this paper cites.
Z3: An efficient smt solver
De Moura, L., and Bjørner, N · 2008
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2020
Earlier work this paper cites.
Automatic testing and improvement of machine translation
Sun, Z., Zhang, J. M., Harman, M., Papadakis, M., and Zhang, L · 2020
Earlier work this paper cites.
Efficient combinatorial optimization for word-level adversarial textual attack
Liu, S., Lu, N., Chen, C., and Tang, K · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
LangChain, Oct. 2022
Chase, H · 2022
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Creswell, A., Shanahan, M., and Higgins, I · 2022
Earlier work this paper cites.
A causal framework to quantify the robustness of mathematical reasoning with language models
Stolfo, A., Jin, Z., Shridhar, K., Schölkopf, B., and Sachan, M · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Earlier work this paper cites.
Interactive joint planning for autonomous vehicles
Chen, Y., Veer, S., Karkus, P., and Pavone, M · 2023
Earlier work this paper cites.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Huang, Y., Shi, J., Li, Y., Fan, C., Wu, S., Zhang, Q., Liu, Y., Zhou, P., Wan, Y., Gong, N. Z., et al · 2023
Earlier work this paper cites.
Benchmarking and explaining large language model-based code generation: A causality-centric approach
Ji, Z., Ma, P., Li, Z., and Wang, S · 2023
Earlier work this paper cites.
Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine
Jiao, W., Wang, W., Huang, J.-t., Wang, X., Shi, S., and Tu, Z · 2023
Earlier work this paper cites.
Les: Locally exploitative sampling for robot path planning
Joshi, S. S., Hutchinson, S., and Tsiotras, P · 2023
Cited alongside, same era.
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukosiute, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., Kadavath, S., Yang, S., Henighan, T., Maxwell, T., Telleen-Lawton, T., Hume, T., Hatfield-Dodds, Z., Kaplan, J., Brauner, J., Bowman, S. R., and Perez, E · 2023
Cited alongside, same era.
CCTEST: Testing and repairing code completion systems
Li, Z., Wang, C., Liu, Z., Wang, H., Wang, S., and Gao, C · 2023
Cited alongside, same era.
Split and merge: Aligning position biases in large language model based evaluators
https://python.langchain.com/docs/modules/agents/agent_types/openai_tools , 2024
Openai tools · 2024
Closest in time.
https://platform.openai.com/docs/guides/function-calling , 2024
Openai’s function callling · 2024
Closest in time.
https://platform.openai.com/docs/assistants/overview?context=with-streaming , 2024
Overview of openai’s assistant · 2024
Closest in time.
https://www.marktechpost.com/2023/10/24/this-ai-research-introduces-rafa-a-principled-artificial-intelligence-framework-for-autonomous-llm-agents-with-provable-sample-efficiency/ , 2024
This ai research introduces ‘rafa’: A principled artificial intelligence framework for autonomous llm agents with provable sample efficiency · 2024
Closest in time.
https://markets.businessinsider.com/news/stocks/unskript-launches-ai-powered-infrastructure-health-intelligence-platform-for-software-teams-1032992108 , 2024
unskript launches ai-powered infrastructure health intelligence platform for software teams · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Z., Wang, C., Ma, P., Wu, D., Li, T., Wang, S., Gao, C., and Liu, Y · 2023
Cited alongside, same era.
Llm+ p: Empowering large language models with optimal planning proficiency
Liu, B., Jiang, Y., Zhang, X., Liu, Q., Zhang, S., Biswas, J., and Stone, P · 2023
Cited alongside, same era.
Can large language models build causal graphs?
Long, S., Schuster, T., Piché, A., de Montreal, U., Research, S., et al · 2023
Cited alongside, same era.
Learning-based near-optimal motion planning for intelligent vehicles with uncertain dynamics
Lu, Y., Zhang, X., Xu, X., and Yao, W · 2023
Cited alongside, same era.
Testing language model agents safely in the wild
Naihin, S., Atkinson, D., Green, M., Hamadi, M., Swift, C., Schonholtz, D., Kalai, A. T., and Bau, D · 2023
Cited alongside, same era.
How susceptible are LLMs to logical fallacies?
Payandeh, A., Pluth, D., Hosier, J., Xiao, X., and Gurbani, V. K · 2023
Cited alongside, same era.
Character-llm: A trainable agent for role-playing
Shao, Y., Li, L., Dai, J., and Qiu, X · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Cited alongside, same era.
Closest in time.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2024
Closest in time.
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., and Yang, Y · 2024
Closest in time.
Combining fine-tuning and LLM-based agents for intuitive smart contract auditing with justifications
Ma, W., Wu, D., Sun, Y., Wang, T., Liu, S., Zhang, J., Xue, Y., and Liu, Y · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Closest in time.
LLM4Vuln: A unified evaluation framework for decoupling and enhancing llms’ vulnerability reasoning
Sun, Y., Wu, D., Xue, Y., Liu, H., Ma, W., Zhang, L., Shi, M., and Liu, Y · 2024
Closest in time.
GPTScan: Detecting logic vulnerabilities in smart contracts by combining GPT with program analysis
Sun, Y., Wu, D., Xue, Y., Liu, H., Wang, H., Xu, Z., Xie, X., and Liu, Y · 2024
Closest in time.
PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., and Kambhampati, S · 2024
Closest in time.
On the planning abilities of large language models-a critical investigation
Valmeekam, K., Marquez, M., Sreedharan, S., and Kambhampati, S · 2024
Closest in time.
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
Wang, S., Wei, Z., Choi, Y., and Ren, X · 2024
Closest in time.
Describe, explain, plan and select: interactive planning with llms enables open-world multi-task agents
Wang, Z., Cai, S., Chen, G., Liu, A., Ma, X. S., and Liang, Y · 2024
Closest in time.
Symbol-llm: Leverage language models for symbolic system in visual human activity reasoning
Wu, X., Li, Y.-L., Sun, J., and Lu, C · 2024
Closest in time.
Satlm: Satisfiability-aided language models using declarative prompting
Ye, X., Chen, Q., Dillig, I., and Durrett, G · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for llm agents
Yuan, T., He, Z., Dong, L., Wang, Y., Zhao, R., Xia, T., Xu, L., Zhou, B., Li, F., Zhang, Z., et al · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Zhan, Q., Liang, Z., Ying, Z., and Kang, D · 2024
Closest in time.