Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have been increasingly employed for (interactive) decision-making, via the development of LLM-based autonomous agents.
Behavior of sequential predictors of binary sequences
Thomas M Cover · 1966
Earlier work this paper cites.
The theory of max-min, with applications
John M Danskin · 1966
Earlier work this paper cites.
Quantal choice analysis: A survey
Daniel L McFadden · 1976
Earlier work this paper cites.
Aggregating strategies
Volodimir G Vovk · 1990
Earlier work this paper cites.
Universal prediction of individual sequences
Meir Feder, Neri Merhav, and Michael Gutman · 1992
Earlier work this paper cites.
Learning mixed equilibria
Drew Fudenberg and David M Kreps · 1993
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K Warmuth · 1994
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard D McKelvey and Thomas R Palfrey · 1995
Earlier work this paper cites.
Worst-case quadratic loss bounds for prediction using linear functions and gradient descent
Nicolo Cesa-Bianchi, Philip M Long, and Manfred K Warmuth · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Predicting how people play games: Reinforcement learning in experimental games with unique, mixed strategy equilibria
Ido Erev and Alvin E Roth · 1998
Earlier work this paper cites.
The theory of learning in games , volume 2
Drew Fudenberg and David K Levine · 1998
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
On the global convergence of stochastic fictitious play
Josef Hofbauer and William H Sandholm · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Strategic learning and its limits
H Peyton Young · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
The topology of the 2x2 games: a new periodic table , volume 3
David Robinson and David Goforth · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
From external to internal regret
Avrim Blum and Yishay Mansour · 2007
Earlier work this paper cites.
Online learning: Theory, algorithms, and applications
Shai Shalev-Shwartz · 2007
Earlier work this paper cites.
A primal-dual perspective of online learning algorithms
Shai Shalev-Shwartz and Yoram Singer · 2007
Earlier work this paper cites.
Regret minimization and the price of total anarchy
Avrim Blum, MohammadTaghi Hajiaghayi, Katrina Ligett, and Aaron Roth · 2008
Earlier work this paper cites.
System reliability theory: Models and statistical methods
Arnljot Hoyland and Marvin Rausand · 2009
Earlier work this paper cites.
Behavioral game theory: Experiments in strategic interaction
Colin F Camerer · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Earlier work this paper cites.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2012
Earlier work this paper cites.
An introduction to order statistics , volume 8
Mohammad Ahsanullah, Valery B Nevzorov, and Mohammad Shakil · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Online linear optimization via smoothing
Jacob Abernethy, Chansoo Lee, Abhinav Sinha, and Ambuj Tewari · 2014
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Earlier work this paper cites.
Fighting bandits with a new kind of smoothness
Jacob Abernethy, Chansoo Lee, and Ambuj Tewari · 2015
Earlier work this paper cites.
Econometrics for learning agents
Denis Nekipelov, Vasilis Syrgkanis, and Eva Tardos · 2015
Earlier work this paper cites.
Intrinsic robustness of the price of anarchy
Tim Roughgarden · 2015
Earlier work this paper cites.
Introduction to online convex optimization
Elad Hazan · 2016
Earlier work this paper cites.
On the properties of the softmax function with application in game theory and reinforcement learning
Bolin Gao and Lacra Pavel · 2017
Earlier work this paper cites.
Beyond the hazard rate: More perturbation algorithms for adversarial multi-armed bandits
Zifan Li and Ambuj Tewari · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Learning in repeated auctions with budgets: Regret minimization and equilibrium
Santiago R Balseiro and Yonatan Gur · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Cited alongside, same era.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu · 2023
Later among the works it cites.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J Horton · 2023
Later among the works it cites.
A latent space theory for emergent abilities in large language models
Hui Jiang · 2023
Later among the works it cites.
Regret minimization via saddle point optimization
Johannes Kirschner, Alireza Bakhtiari, Kushagra Chandak, Volodymyr Tkachuk, and Csaba Szepesvari · 2023
Later among the works it cites.
In-context reinforcement learning with algorithm distillation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Başar · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Near-optimal no-regret learning in general games
Constantinos Daskalakis, Maxwell Fishelson, and Noah Golowich · 2021
Cited alongside, same era.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Cited alongside, same era.
Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Cited alongside, same era.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, et al · 2022
Cited alongside, same era.
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, et al · 2023
Later among the works it cites.
Supervised pretraining can learn in-context reinforcement learning
Jonathan N Lee, Annie Xie, Aldo Pacchiano, Yash Chandak, Chelsea Finn, Ofir Nachum, and Emma Brunskill · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi · 2023
Later among the works it cites.
Reason for future, act for now: A principled architecture for autonomous llm agents
Zhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo, Shuqi Ke, Boyi Liu, and Zhaoran Wang · 2023
Later among the works it cites.
LLM Engine, 2023
LLM Engine · 2023
Later among the works it cites.
Strategic behavior of large language models: Game structure vs. contextual framing
Nunzio Lorè and Babak Heydari · 2023
Later among the works it cites.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind Mahankali, Tatsunori B Hashimoto, and Tengyu Ma · 2023
Later among the works it cites.
Welfare diplomacy: Benchmarking language model cooperation
Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer, Lewis Hammond, Alan Chan, and Jesse Clifton · 2023
Later among the works it cites.
Gpt-4 technical report
Openai · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Later among the works it cites.
Communicative agents for software development
Chen Qian, Xin Cong, Cheng Yang, Weize Chen, Yusheng Su, Juyuan Xu, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Peer: A collaborative language model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel · 2023
Later among the works it cites.
Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Autogpt, 2023
Significant Gravitas · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2023
Later among the works it cites.
Math agents: Computational infrastructure, mathematical embedding, and genomics
Melanie Swan, Takashi Kido, Eric Roland, and Renato P dos Santos · 2023
Later among the works it cites.
Can large language models play text games well? current state-of-the-art and open questions
Chen Feng Tsai, Xiaochen Zhou, Sierra S Liu, Jing Li, Mo Yu, and Hongyuan Mei · 2023
Later among the works it cites.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Large language models are implicitly topic models: Explaining and finding good demonstrations for in-context learning
Xinyi Wang, Wanrong Zhu, and William Yang Wang · 2023
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Later among the works it cites.
Examining inter-consistency of large language models collaboration: An in-depth analysis via debate
Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, and Bing Qin · 2023
Later among the works it cites.
Competeai: Understanding the competition behaviors in large language model-based agents
Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, and Xing Xie · 2023
Later among the works it cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2024
Closest in time.
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents
Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chen Qian, Chi-Min Chan, Yujia Qin, Yaxi Lu, Ruobing Xie, et al · 2024
Closest in time.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, et al · 2024
Closest in time.
Can large language models explore in-context?
Akshay Krishnamurthy, Keegan Harris, Dylan J Foster, Cyril Zhang, and Aleksandrs Slivkins · 2024
Closest in time.
Transformers as decision makers: Provable in-context reinforcement learning via supervised pretraining
Licong Lin, Yu Bai, and Song Mei · 2024
Closest in time.
Beyond numeric awards: In-context dueling bandits with llm agents
Fanzeng Xia, Hao Liu, Yisong Yue, and Tongxin Li · 2024
Closest in time.
Building cooperative embodied agents modularly with large language models
Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B Tenenbaum, Tianmin Shu, and Chuang Gan · 2024
Closest in time.
Chanwoo Park, Seungju Han, Xingzhi Guo, Asuman Ozdaglar, Kaiqing Zhang, and Joo-Kyung Kim · 2025
Closest in time.