Fetching the paper…
Reading the bibliography…
Recent advancements in reinforcement learning (RL) for large language models (LLMs), exemplified by DeepSeek R1, have shown that even a simple question-answering task can substantially improve an LLM's reasoning capabilities.
A context-aware natural language generator for dialogue systems
Ondřej Dušek and Filip Jurčíček · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Context-aware language modeling for goal-oriented dialogue systems, 2022
Charlie Snell, Mengjiao Yang, Justin Fu, Yi Su, and Sergey Levine · 2022
Earlier work this paper cites.
Yanzhao Zheng, Haibin Wang, Baohua Dong, Xingjun Wang, and Changshan Li · 2022
Earlier work this paper cites.
Thinker: Learning to plan and act
Stephen Chung, Ivan Anokhin, and David Krueger · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Earlier work this paper cites.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen · 2023
Cited alongside, same era.
Active retrieval augmented generation
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Cited alongside, same era.
Learning to reason with llms
OpenAI · 2024
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al · 2024
Cited alongside, same era.
Archer: Training language model agents via hierarchical multi-turn rl
Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar · 2024
Direct multi-turn preference optimization for language agents
Wentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang, and Fuli Feng · 2024
Later among the works it cites.
When can llms actually correct their own mistakes? a critical survey of self-correction of llms
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang · 2024
Later among the works it cites.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Multi-turn code generation through single-step rewards, 2025
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M Rush, Wenting Zhao, and Sanjiban Choudhury · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-turn reinforcement learning from preference human feedback
Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, et al · 2024
Cited alongside, same era.
Closest in time.
7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Weihao Zeng, Yuzhen Huang, Wei Liu, Keqing He, Qian Liu, Zejun Ma, and Junxian He · 2025
Closest in time.