Fetching the paper…
Reading the bibliography…
The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I. Levenshtein. 1966 · 1966
Earlier work this paper cites.
Lateral Thinking Puzzlers
Paul Sloane. 1992 · 1992
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Mastermind is NP-complete
Jeff Stuckman and Guo-Qiang Zhang. 2005 · 2005
Earlier work this paper cites.
Mathematics of Sudoku I
Bertram Felgenhauer and Frazer Jarvis. 2006 · 2006
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Earlier work this paper cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020 · 2020
Cited alongside, same era.
LangChain - Building applications with LLMs through composability
Harrison Chase. 2022 · 2022
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Cited alongside, same era.
Clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational Agents
GAIA: a benchmark for General AI Assistants
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Generative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Later among the works it cites.
Gorilla: Large Language Model Connected with Massive APIs
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023 · 2023
Later among the works it cites.
A Survey on Large Language Model based Autonomous Agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2023 · 2023
Later among the works it cites.
LLM-powered Autonomous Agents
Lilian Weng. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kranti Chalamalasetti, Jana Götze, Sherzod Hakimov, Brielen Madureira, Philipp Sadler, and David Schlangen. 2023 · 2023
Cited alongside, same era.
Plotting Progress in AI
Douwe Kiela, Tristan Thrush, Kawin Ethayarajh, and Amanpreet Singh. 2023 · 2023
Cited alongside, same era.
AgentBench: Evaluating LLMs as Agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al. 2023 · 2023
Cited alongside, same era.
Assistants API
OpenAI. 2023a
Cited in the paper.
OpenAI. 2023b
Cited in the paper.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020a
Cited in the paper.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020b
Cited in the paper.
Later among the works it cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
ToolQA: A Dataset for LLM Question Answering with External Tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2023 · 2023
Later among the works it cites.