Fetching the paper…
Reading the bibliography…
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild.
“Measuring Coding Challenge Competence With APPS”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“Measuring Mathematical Problem Solving With the MATH Dataset”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“Are we learning yet? a meta review of evaluation failures across machine learning”
Thomas Liao, Rohan Taori, Inioluwa Raji and Ludwig Schmidt · 2021
Earlier work this paper cites.
“Holistic evaluation of language models”
Percy Liang et al · 2022
Earlier work this paper cites.
“Chain of thought prompting elicits reasoning in large language models”
Jason Wei et al · 2022
Earlier work this paper cites.
“Mind2Web: Towards a Generalist Agent for the Web” type: article, 2023
Xiang Deng et al · 2023
Earlier work this paper cites.
“Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text”
Sebastian Gehrmann, Elizabeth Clark and Thibault Sellam · 2023
Cited alongside, same era.
“BabyAGI” original-date: 2023-04-03T00:40:27Z, 2023
Yohei Nakajima · 2023
Cited alongside, same era.
“ChatGPT plugins”, 2023
OpenAI · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
“Auto-GPT: An Autonomous GPT-4 Experiment” original-date: 2023-03-16T09:21:07Z
Toran Richards · 2023
Cited alongside, same era.
“Are Emergent Abilities of Large Language Models a Mirage?” type: article, 2023
“Reflexion: Language Agents with Verbal Reinforcement Learning” type: article, 2023
Noah Shinn et al · 2023
Closest in time.
Aarohi Srivastava et al · 2023
Closest in time.
“SuperAGI” original-date: 2023-05-13T08:55:24Z, 2023
SuperAGI · 2023
Closest in time.
“Voyager: An Open-Ended Embodied Agent with Large Language Models” type: article, 2023
Guanzhi Wang et al · 2023
Closest in time.
“ReAct: Synergizing Reasoning and Acting in Language Models”
Shunyu Yao et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Cited alongside, same era.
“Model evaluation for extreme risks”
Toby Shevlane et al · 2023
Cited alongside, same era.
“WebArena: A Realistic Web Environment for Building Autonomous Agents”, 2023
Shuyan Zhou et al · 2023
Closest in time.