Fetching the paper…
Reading the bibliography…
Artificial intelligence progresses towards the "Era of Experience," where agents are expected to learn from continuous, grounded interaction.
OpenSpiel: A framework for reinforcement learning in games
M. Lanctot, E. Lockhart, J.-B. Lespiau, V. Zambaldi, S. Upadhyay, J. Pérolat, S. Srinivasan, F. Timbers, K. Tuyls, S. Omidshafiei, D. Hennes, D. Morrill, P. Muller, T. Ewalds, R. Faulkner, J. Kramár, B. D. Vylder, B. Saeta, J. Bradbury, D. Ding, S. Borgeaud, M. Lai, J. Schrittwieser, T. Anthony, E. Hughes, I. Danihelka, and J. Ryan-Davis · 1908
Earlier work this paper cites.
Dynamic programming and modern control theory , volume 81
R. Bellman, R. E. Kalaba, et al · 1965
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Solving breakthrough with race patterns and job-level proof number search
A. Saffidine, N. Jouandeau, and T. Cazenave · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. P. van Hasselt, and D. Silver · 2017
Earlier work this paper cites.
Improving robot controller transparency through autonomous policy explanation
B. Hayes and J. A. Shah · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
S. Sreedharan, U. Soni, M. Verma, S. Srivastava, and S. Kambhampati · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Cited alongside, same era.
Lmrl gym: Benchmarks for multi-turn reinforcement learning with language models
M. Abdulhai, I. White, C. Snell, C. Sun, J. Hong, Y. Zhai, K. Xu, and S. Levine · 2023
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Cited alongside, same era.
Llf-bench: Benchmark for interactive learning from language feedback
Benchmarking large language models for news summarization
T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto · 2023
Later among the works it cites.
Trace is the next autodiff: Generative optimization with rich feedback, execution traces, and llms
C.-A. Cheng, A. Nie, and A. Swaminathan · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
Llm-based nlg evaluation: Current status and challenges
M. Gao, X. Hu, J. Ruan, X. Pu, and X. Wan · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-A. Cheng, A. Kolobov, D. Misra, A. Nie, and A. Swaminathan · 2023
Cited alongside, same era.
Pangu-agent: A fine-tunable generalist agent with structured reasoning
F. Christianos, G. Papoudakis, M. Zimmer, T. Coste, Z. Wu, J. Chen, K. Khandelwal, J. Doran, X. Feng, J. Liu, et al · 2023
Cited alongside, same era.
State2explanation: Concept-based explanations to benefit agent learning and user understanding
D. Das, S. Chernova, and B. Kim · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu · 2023
Cited alongside, same era.
Spatial-temporal concept based explanation of 3d convnets
Y. Ji, Y. Wang, and J. Kato · 2023
Cited alongside, same era.
Tigerscore: Towards building explainable metric for all text generation tasks
D. Jiang, Y. Li, G. Zhang, W. Huang, B. Y. Lin, and W. Chen · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Generative judge for evaluating alignment
J. Li, S. Sun, W. Yuan, R.-Z. Fan, H. Zhao, and P. Liu · 2023
Cited alongside, same era.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Closest in time.
gym-tictactoe: A tic tac toe environment for openai gym, 2024
K. J. Ju · 2024
Closest in time.
D. Mahan, D. Van Phung, R. Rafailov, C. Blagden, N. Lile, L. Castricato, J.-P. Fränken, C. Finn, and A. Albalak · 2024
Closest in time.
Agent q: Advanced reasoning and learning for autonomous ai agents
P. Putta, E. Mills, N. Garg, S. Motwani, C. Finn, D. Garg, and R. Rafailov · 2024
Closest in time.
Textgrad: Automatic" differentiation" via text
M. Yuksekgonul, F. Bianchi, J. Boen, S. Liu, Z. Huang, C. Guestrin, and J. Zou · 2024
Closest in time.
Generative verifiers: Reward modeling as next-token prediction
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2024
Closest in time.
Policy improvement using language feedback models
V. Zhong, D. Misra, X. Yuan, and M.-A. Côté · 2024
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Search-r1: Training llms to reason and leverage search engines with reinforcement learning
B. Jin, H. Zeng, Z. Yue, J. Yoon, S. Arik, D. Wang, H. Zamani, and J. Han · 2025
Closest in time.
Understanding r1-zero-like training: A critical perspective
Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin · 2025
Closest in time.
Rethinking reflection in pre-training
D. J. Shah, P. Rushton, S. Singla, M. Parmar, K. Smith, Y. Vanjani, A. Vaswani, A. Chaluvaraju, A. Hojel, A. Ma, et al · 2025
Closest in time.
Welcome to the era of experience
D. Silver and R. S. Sutton · 2025
Closest in time.
Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning
Z. Wang, K. Wang, Q. Wang, P. Zhang, L. Li, Z. Yang, K. Yu, M. N. Nguyen, L. Liu, E. Gottlieb, et al · 2025
Closest in time.