Fetching the paper…
Reading the bibliography…
Large language models (LLMs) demonstrate impressive reasoning abilities, but translating reasoning into actions in the real world remains challenging.
Language models are few-shot learners
Brown, T · 1901
Earlier work this paper cites.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
Dong, K · 1901
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Bickel, P. J · 1981
Earlier work this paper cites.
On the asymptotic accuracy of efron’s bootstrap
Singh, K · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans
Efron, B · 1982
Earlier work this paper cites.
Approximate bayesian inference with the weighted likelihood bootstrap
Newton, M. A · 1994
Earlier work this paper cites.
Bootstrap methods and their application
Davison, A. C · 1997
Earlier work this paper cites.
Survey of numerical methods for trajectory optimization
Betts, J. T · 1998
Earlier work this paper cites.
Model predictive control: past, present and future
Morari, M · 1999
Earlier work this paper cites.
Tutorial overview of model predictive control
Rawlings, J. B · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Bayesian model selection and model averaging
Wasserman, L · 2000
Earlier work this paper cites.
Combinatorial games: Tic-Tac-Toe theory
Beck, J · 2008
Earlier work this paper cites.
Analysis of a classification-based policy iteration algorithm
Lazaric, A · 2010
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Osband, I · 2013
Earlier work this paper cites.
Bayesian optimal control of smoothly parameterized systems
Abbasi-Yadkori, Y · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I · 2016
Earlier work this paper cites.
An information-theoretic analysis of Thompson sampling
Russo, D · 2016
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Hafner, D · 2019
Earlier work this paper cites.
Bootstrapping upper confidence bound
Hao, B · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M · 2019
Earlier work this paper cites.
Information-theoretic confidence bounds for reinforcement learning
Lu, X · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A · 2019
Earlier work this paper cites.
Sample-optimal parametric q q -learning using linearly additive features
Yang, L · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Cited alongside, same era.
A theoretical analysis of deep q-learning
Fan, J · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Cited alongside, same era.
Planning to explore via self-supervised world models
Sekar, R · 2020
Cited alongside, same era.
An analysis of attention via the lens of exchangeability and latent variable models
Zhang, Y · 2022
Later among the works it cites.
Optimistic exploration with learned features provably solves markov decision processes with neural dynamics
Zheng, S · 2022
Later among the works it cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Zhong, H · 2022
Later among the works it cites.
A mechanism for sample-efficient in-context learning for sparse retrieval tasks
Abernethy, J · 2023
Closest in time.
Large language models as tool makers
Cai, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Model-free reinforcement learning in infinite-horizon average-reward Markov decision processes
Wei, C.-Y · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L · 2020
Cited alongside, same era.
Generative adversarial imitation learning with neural network parameterization: Global optimality and convergence rate
Zhang, Y · 2020
Cited alongside, same era.
Exponential tail bounds for chisquared random variables
Ghosh, M · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit Bayesian inference
Xie, S. M · 2021
Cited alongside, same era.
Do as I can, not as I say: Grounding language in robotic affordances
Ahn, M · 2022
Cited alongside, same era.
Closest in time.
Large language models are zero-shot time series forecasters
Gruver, N · 2023
Closest in time.
Reasoning with language model is planning with world model
Hao, S · 2023
Closest in time.
A latent space theory for emergent abilities in large language models
Jiang, H · 2023
Closest in time.
Language models can solve computer tasks
Kim, G · 2023
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Lee, J. N · 2023
Closest in time.
Transformers as algorithms: Generalization and implicit model selection in in-context learning
Li, Y · 2023
Closest in time.
Chameleon: Plug-and-play compositional reasoning with large language models
Lu, P · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
REFINER: Reasoning feedback on intermediate representations
Paul, D · 2023
Closest in time.
Algorithm of thoughts: Enhancing exploration of ideas in large language models
Sel, B · 2023
Closest in time.
HuggingGPT: Solving AI tasks with ChatGPT and its friends in HuggingFace
Shen, Y · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N · 2023
Closest in time.
Reinforcement learning in the era of llms: What is essential? what is needed? an rl perspective on rlhf, prompting, and beyond
Sun, H · 2023
Closest in time.
Function vectors in large language models
Todd, E · 2023
Closest in time.
LLaMa: Open and efficient foundation language models
Touvron, H · 2023
Closest in time.
Transformers learn in-context by gradient descent
Von Oswald, J · 2023
Closest in time.
The learnability of in-context learning
Wies, N · 2023
Closest in time.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L · 2023
Closest in time.
Can large language models play games? a case study of a self-play approach
Guo, H · 2024
Closest in time.
Maximize to explore: One objective function fusing estimation, planning, and exploration
Liu, Z · 2024
Closest in time.
Retrieval-augmented thought process as sequential decision making
Pouplin, T · 2024
Closest in time.
The learnability of in-context learning
Wies, N · 2024
Closest in time.
How can llm guide rl? a value-based approach
Zhang, S · 2024
Closest in time.