Fetching the paper…
Reading the bibliography…
The success of Large Language Models (LLMs) has sparked interest in various agentic applications.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 1901
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 1901
Earlier work this paper cites.
Evolutionary principles in self-referential learning. on learning now to learn: The meta-meta-meta…-hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
P.-Y. Oudeyer, F. Kaplan, and V. V. Hafner · 2007
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
S. Still and D. Precup · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
The movielens datasets: History and context
F. M. Harper and J. A. Konstan · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Meta-learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Earlier work this paper cites.
A simple neural attentive meta-learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
N. Shazeer and M. Stern · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Rudder: Return decomposition for delayed rewards
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2019
Earlier work this paper cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Earlier work this paper cites.
Meta-learning with warped gradient descent
S. Flennerhag, A. A. Rusu, R. Pascanu, F. Visin, H. Yin, and R. Hadsell · 2019
Earlier work this paper cites.
Improving generalization in meta reinforcement learning using learned objectives
L. Kirsch, S. van Steenkiste, and J. Schmidhuber · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Introduction to multi-armed bandits
A. Slivkins et al · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
M. G. Bellemare, S. Candido, P. S. Castro, J. Gong, M. C. Machado, S. Moitra, S. S. Ponda, and Z. Wang · 2020
Cited alongside, same era.
Bandit algorithms
Towards general-purpose in-context learning agents
L. Kirsch, J. Harrison, C. Freeman, J. Sohl-Dickstein, and J. Schmidhuber · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback
M. Klissarov, P. D’Oro, S. Sodhani, R. Raileanu, P.-L. Bacon, P. Vincent, A. Zhang, and M. Henaff · 2023
Later among the works it cites.
Gaia: a benchmark for general ai assistants
G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom · 2023
Later among the works it cites.
Large language models as general pattern machines
S. Mirchandani, F. Xia, P. Florence, B. Ichter, D. Driess, M. G. Arenas, K. Rao, D. Sadigh, and A. Zeng · 2023
Later among the works it cites.
Generalization to new sequential decision making tasks with in-context learning, 2023
S. C. Raparthy, E. Hambro, R. Kirk, M. Henaff, and R. Raileanu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
R. Raileanu and T. Rocktäschel · 2020
Cited alongside, same era.
Fighting copycat agents in behavioral cloning from observation histories
C. Wen, J. Lin, T. Darrell, D. Jayaraman, and Y. Gao · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Cited alongside, same era.
Is curiosity all you need? on the utility of emergent behaviours from curious exploration
O. Groth, M. Wulfmeier, G. Vezzani, V. Dasagi, T. Hertweck, R. Hafner, N. Heess, and M. Riedmiller · 2021
Cited alongside, same era.
Benchmarking the spectrum of agent capabilities
D. Hafner · 2021
Cited alongside, same era.
Learning to modulate pre-trained models in rl
T. Schmied, M. Hofmarcher, F. Paischer, R. Pascanu, and S. Hochreiter · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning.(2023)
N. Shinn, F. Cassano, B. Labash, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, et al · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, A. Üstün, and S. Hooker · 2024
Later among the works it cites.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, et al · 2024
Later among the works it cites.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Later among the works it cites.
Agent ai: Surveying the horizons of multimodal interaction
Z. Durante, Q. Huang, N. Wake, R. Gong, J. S. Park, B. Sarkar, R. Taori, Y. Noda, D. Terzopoulos, Y. Choi, et al · 2024
Later among the works it cites.
Webvoyager: Building an end-to-end web agent with large multimodal models
H. He, W. Yao, K. Ma, W. Yu, Y. Dai, H. Zhang, Z. Lan, and D. Yu · 2024
Later among the works it cites.
Maestromotif: Skill design from artificial intelligence feedback
M. Klissarov, M. Henaff, R. Raileanu, S. Sodhani, P. Vincent, A. Zhang, P.-L. Bacon, D. Precup, M. C. Machado, and P. D’Oro · 2024
Later among the works it cites.
Can large language models explore in-context?
A. Krishnamurthy, K. Harris, D. J. Foster, C. Zhang, and A. Slivkins · 2024
Later among the works it cites.
Training language models to self-correct via reinforcement learning
A. Kumar, V. Zhuang, R. Agarwal, Y. Su, J. D. Co-Reyes, A. Singh, K. Baumli, S. Iqbal, C. Bishop, R. Roelofs, et al · 2024
Later among the works it cites.
Intelligent go-explore: Standing on the shoulders of giant foundation models
C. Lu, S. Hu, and J. Clune · 2024
Later among the works it cites.
Llms are in-context reinforcement learners
G. Monea, A. Bosselut, K. Brantley, and Y. Artzi · 2024
Later among the works it cites.
Evolve: Evaluating and optimizing llms for exploration
A. Nie, Y. Su, B. Chang, J. N. Lee, E. H. Chi, Q. V. Le, and M. Chen · 2024
Later among the works it cites.
Balrog: Benchmarking agentic llm and vlm reasoning on games
D. Paglieri, B. Cupiał, S. Coward, U. Piterbarg, M. Wolczyk, A. Khan, E. Pignatelli, Ł. Kuciński, L. Pinto, R. Fergus, et al · 2024
Later among the works it cites.
Group robust preference optimization in reward-free rlhf
S. S. Ramesh, Y. Hu, I. Chaimalas, V. Mehta, P. G. Sessa, H. B. Ammar, and I. Bogunovic · 2024
Later among the works it cites.
Lmact: A benchmark for in-context imitation learning with long multimodal demonstrations
A. Ruoss, F. Pardo, H. Chan, B. Li, V. Mnih, and T. Genewein · 2024
Later among the works it cites.
Capabilities of gemini models in medicine
K. Saab, T. Tu, W.-H. Weng, R. Tanno, D. Stutz, E. Wulczyn, F. Zhang, T. Strother, C. Park, E. Vedadi, et al · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Later among the works it cites.
Imitating language via scalable inverse reinforcement learning
M. Wulfmeier, M. Bloesch, N. Vieillard, A. Ahuja, J. Bornschein, S. Huang, A. Sokolov, M. Barnes, G. Desjardins, A. Bewley, S. M. E. Bechtle, J. T. Springenberg, N. Momchev, O. Bachem, M. Geist, and M. Riedmiller · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman · 2024
Later among the works it cites.
In-context principle learning from mistakes
T. Zhang, A. Madaan, L. Gao, S. Zheng, S. Mishra, Y. Yang, N. Tandon, and U. Alon · 2024
Later among the works it cites.
xlstm: Extended long short-term memory
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto · 2025
Closest in time.
Learning robot soccer from egocentric vision with deep reinforcement learning
D. Tirumala, M. Wulfmeier, B. Moran, S. Huang, J. Humplik, G. Lever, T. Haarnoja, L. Hasenclever, A. Byravan, N. Batchelor, N. sreendra, K. Patel, M. Gwira, F. Nori, M. Riedmiller, and N. Heess · 2025
Closest in time.
Critique fine-tuning: Learning to critique is more effective than learning to imitate
Y. Wang, X. Yue, and W. Chen · 2025
Closest in time.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
S. Zhai, H. Bai, Z. Lin, J. Pan, P. Tong, Y. Zhou, A. Suhr, S. Xie, Y. LeCun, Y. Ma, et al · 2025
Closest in time.