Fetching the paper…
Reading the bibliography…
A burgeoning area within reinforcement learning (RL) is the design of sequential decision-making agents centered around large language models (LLMs).
On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
William R Thompson · 1933
Earlier work this paper cites.
A Markovian Decision Process
Richard Bellman · 1957
Earlier work this paper cites.
On Adaptive Control Processes
Richard Bellman and Robert Kalaba · 1959
Earlier work this paper cites.
Dynamic Programming and Markov Processes
Ronald A Howard · 1960
Earlier work this paper cites.
Bandit Processes and Dynamic Allocation Indices
John Gittins · 1979
Earlier work this paper cites.
Asymptotically Efficient Adaptive Allocation Rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Self-Improving Reactive Agents Based on Reinforcement learning, Planning and Teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Q Q -Learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
A Bayesian Framework for Reinforcement Learning
Malcolm JA Strens · 2000
Earlier work this paper cites.
Finite-Time Analysis of the Multiarmed Bandit Problem
P Auer, Paul Fischer, and N Cesa-Bianchi · 2002
Earlier work this paper cites.
R-MAX - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Optimal Learning: Computational Procedures for Bayes-Adaptive Markov Decision Processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
Near-Optimal Reinforcement Learning in Polynomial Time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
An Analysis of Model-Based Interval Estimation for Markov Decision Processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Near-Optimal Regret Bounds for Reinforcement Learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
Aleatory or Epistemic? Does it Matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
Reinforcement Learning in Finite MDPs: PAC Analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Near-Optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
An Empirical Evaluation of Thompson Sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
The k k -Armed Dueling Bandits Problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims · 2012
Earlier work this paper cites.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Bayesian Optimal Control of Smoothly Parameterized Systems: The Lazy Posterior Sampling Algorithm
Yasin Abbasi-Yadkori and Csaba Szepesvari · 2014
Earlier work this paper cites.
Model-Based Reinforcement Learning and the Eluder Dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Learning to Optimize via Posterior Sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Sample Complexity of Episodic Fixed-Horizon Reinforcement Learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
Contextual Dueling Bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Bayesian Reinforcement Learning: A Survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Earlier work this paper cites.
The Dependence of Effective Planning Horizon on Model Accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Earlier work this paper cites.
Dropout as a Bayesian Epproximation: Representing Model Uncertainty in Deep Learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Posterior Sampling for Reinforcement Learning Without Episodes
Ian Osband and Benjamin Van Roy · 2016
Earlier work this paper cites.
An Information-Theoretic Analysis of Thompson Sampling
Daniel Russo and Benjamin Van Roy · 2016
Earlier work this paper cites.
Optimistic Posterior Sampling for Reinforcement Learning: Worst-Case Regret Bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Minimax Regret Bounds for Reinforcement Learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Ensemble Sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Cited alongside, same era.
Why is Posterior Sampling Better than Optimism for Reinforcement Learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
Learning Unknown Markov Decision Processes: A Thompson Sampling Approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Cited alongside, same era.
Mitigating Planner Overfitting in Model-Based Reinforcement Learning
Dilip Arumugam, David Abel, Kavosh Asadi, Nakul Gopalan, Christopher Grimm, Jun Ki Lee, Lucas Lehnert, and Michael L Littman · 2018
Large Language Models Can Implement Policy Iteration
Ethan Brooks, Logan Walls, Richard L Lewis, and Satinder Singh · 2023
Later among the works it cites.
Meta-In-Context Learning in Large Language Models
Julian Coda-Forno, Marcel Binz, Zeynep Akata, Matt Botvinick, Jane Wang, and Eric Schulz · 2023
Later among the works it cites.
Social Contract AI: Aligning AI Assistants with Implicit Group Norms
Jan-Philipp Fränken, Sam Kwok, Peixuan Ye, Kanishk Gandhi, Dilip Arumugam, Jared Moore, Alex Tamkin, Tobias Gerstenberg, and Noah D Goodman · 2023
Later among the works it cites.
Meta-Prompt: A Simple Self-Improving Language Agent
Noah Goodman · 2023
Later among the works it cites.
Langevin Thompson Sampling with Logarithmic Communication: Bandits and Reinforcement Learning
Amin Karbasi, Nikki Lijing Kuang, Yian Ma, and Siddharth Mitra · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is Q Q -Learning Provably Efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
Cited alongside, same era.
The Uncertainty Bellman Equation and Exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2018
Cited alongside, same era.
Randomized Prior Functions for Deep Reinforcement Learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Learning to Optimize via Information-Directed Sampling
Daniel Russo and Benjamin Van Roy · 2018
Cited alongside, same era.
A Tutorial on Thompson Sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Cited alongside, same era.
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh · 2023
Later among the works it cites.
Zhihan Liu, Hao Hu, Shenao Zhang, Hongyi Guo, Shuqi Ke, Boyi Liu, and Zhaoran Wang · 2023
Later among the works it cites.
Reinforcement Learning, Bit by Bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2023
Later among the works it cites.
Approximate Thompson Sampling via Epistemic Neural Networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy · 2023
Later among the works it cites.
The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning
Sarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan Murphy, and Finale Doshi-Velez · 2023
Later among the works it cites.
Posterior Sampling for Deep Reinforcement Learning
Remo Sasso, Michelangelo Conserva, and Paulo Rauber · 2023
Later among the works it cites.
Probabilistic Inference in Reinforcement Learning Done Right
Jean Tarbouriech, Tor Lattimore, and Brendan O’Donoghue · 2023
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Shattering the Agent-Environment Interface for Fine-Tuning Inclusive Language Models
Wanqiao Xu, Shi Dong, Dilip Arumugam, and Benjamin Van Roy · 2023
Later among the works it cites.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Yufeng Zhang, Fengzhuo Zhang, Zhuoran Yang, and Zhaoran Wang · 2023
Later among the works it cites.
CogBench: A Large Language Model Walks into a Psychology Lab
Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz · 2024
Later among the works it cites.
In-Context Exploration-Exploitation for Reinforcement Learning
Zhenwen Dai, Federico Tomasi, and Sina Ghiassian · 2024
Later among the works it cites.
Efficient Exploration for LLMs
Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, and Benjamin Van Roy · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte Carlo
Haque Ishfaq, Qingfeng Lan, Pan Xu, A Rupam Mahmood, Doina Precup, Anima Anandkumar, and Kamyar Azizzadenesheli · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Later among the works it cites.
Isoperimetry is All We Need: Langevin Posterior Sampling for RL with Sublinear Regret
Emilio Jorge, Christos Dimitrakakis, and Debabrota Basu · 2024
Later among the works it cites.
Can Foundation Models actively Gather Information in Interactive Environments to Test Hypotheses?
Nan Rosemary Ke, Danny P Sawyer, Hubert Soyer, Martin Engelcke, David P Reichert, Drew A Hudson, John Reid, Alexander Lerchner, Danilo Jimenez Rezende, Timothy P Lillicrap, Michael Mozer, and Jane X Wang · 2024
Later among the works it cites.
Can Large Language Models Explore In-Context?
Akshay Krishnamurthy, Keegan Harris, Dylan J Foster, Cyril Zhang, and Aleksandrs Slivkins · 2024
Later among the works it cites.
Embers of Autoregression Show how Large Language Models are Shaped by the Problem They are Trained to Solve
R Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D Hardy, and Thomas L Griffiths · 2024
Later among the works it cites.
LLMs Are In-Context Reinforcement Learners
Giovanni Monea, Antoine Bosselut, Kianté Brantley, and Yoav Artzi · 2024
Later among the works it cites.
EVOLvE: Evaluating and Optimizing LLMs For Exploration
Allen Nie, Yi Su, Bo Chang, Jonathan N Lee, Ed H Chi, Quoc V Le, and Minmin Chen · 2024
Later among the works it cites.
Reflexion: Language Agents with Verbal Reinforcement Learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2024
Later among the works it cites.
Posterior Sampling for Continuing Environments
Wanqiao Xu, Shi Dong, and Benjamin Van Roy · 2024
Later among the works it cites.
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
Qinqing Zheng, Mikael Henaff, Amy Zhang, Aditya Grover, and Brandon Amos · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
Martin Klissarov, R Devon Hjelm, Alexander T Toshev, and Bogdan Mazoure · 2025
Closest in time.
Welcome to the Era of Experience
David Silver and Richard S Sutton · 2025
Closest in time.
Training a Generally Curious Agent
Fahim Tajwar, Yiding Jiang, Abitha Thankaraj, Sumaita Sadia Rahman, J Zico Kolter, Jeff Schneider, and Russ Salakhutdinov · 2025
Closest in time.
Efficient Reinforcement Learning with Large Language Model Priors
Xue Yan, Yan Song, Xidong Feng, Mengyue Yang, Haifeng Zhang, Haitham Bou Ammar, and Jun Wang · 2025
Closest in time.