Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), when enhanced through reasoning-oriented post-training, evolve into powerful Large Reasoning Models (LRMs).
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
Natural Questions: A Benchmark for Question Answering Research
Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; et al. 2019 · 2019
Earlier work this paper cites.
Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Ho, X.; Nguyen, A.-K. D.; Sugawara, S.; and Aizawa, A. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
LogiQA: a challenge dataset for machine reading comprehension with logical reasoning
Liu, J.; Cui, L.; Liu, H.; Huang, D.; Wang, Y.; and Zhang, Y. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
MuSiQue: Multihop Questions via Single-hop Question Composition
Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2022 · 2022
Earlier work this paper cites.
Text embeddings by weakly-supervised contrastive pre-training
Wang, L.; Yang, N.; Huang, X.; Jiao, B.; Yang, L.; Jiang, D.; Majumder, R.; and Wei, F. 2022 · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Zelikman, E.; Wu, Y.; Mu, J.; and Goodman, N. 2022 · 2022
Earlier work this paper cites.
Shortcut learning of large language models in natural language understanding
Du, M.; He, F.; Zou, N.; Tao, D.; and Hu, X. 2023 · 2023
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Gou, Z.; Shao, Z.; Gong, Y.; Shen, Y.; Yang, Y.; Huang, M.; Duan, N.; and Chen, W. 2023 · 2023
Earlier work this paper cites.
Training chain-of-thought via latent-variable inference
Hoffman, M. D.; Phan, D.; Dohan, D.; Douglas, S.; Le, T. A.; Parisi, A.; Sountsov, P.; Sutton, C.; Vikram, S.; and Saurous, R. A. 2023 · 2023
Earlier work this paper cites.
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; and Li, Y. 2023 · 2023
Earlier work this paper cites.
Measuring and Narrowing the Compositionality Gap in Language Models
Press, O.; Zhang, M.; Min, S.; Schmidt, L.; Smith, N. A.; and Lewis, M. 2023 · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023 · 2023
Earlier work this paper cites.
Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy
Shao, Z.; Gong, Y.; Shen, Y.; Huang, M.; Duan, N.; and Chen, W. 2023 · 2023
Earlier work this paper cites.
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2023 · 2023
Earlier work this paper cites.
Wei, Y.; Su, Y.; Ma, H.; Yu, X.; Lei, F.; Zhang, Y.; Zhao, J.; and Liu, K. 2023 · 2023
Cited alongside, same era.
Gpt4tools: Teaching large language model to use tools via self-instruction
Yang, R.; Song, L.; Li, Y.; Zhao, S.; Ge, Y.; Li, X.; and Shan, Y. 2023 · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models
Zhou, J.; Lu, T.; Mishra, S.; Brahma, S.; Basu, S.; Luan, Y.; Zhou, D.; and Hou, L. 2023 · 2023
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering
Adlakha, V.; BehnamGhader, P.; Lu, X. H.; Meade, N.; and Reddy, S. 2024 · 2024
Scaling reasoning, losing control: Evaluating instruction following in large reasoning models
Fu, T.; Gu, J.; Li, Y.; Qu, X.; and Cheng, Y. 2025 · 2025
Closest in time.
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Gehring, J.; Zheng, K.; Copet, J.; Mella, V.; Cohen, T.; and Synnaeve, G. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025 · 2025
Closest in time.
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team
Hosain, M. T.; Rahman, S.; Morol, M. K.; and Parvez, M. R. 2025 · 2025
Closest in time.
Reinforced Internal-External Knowledge Synergistic Reasoning for Efficient Adaptive Search Agent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MATHSENSEI: a tool-augmented large language model for mathematical reasoning
Das, D.; Banerjee, D.; Aditya, S.; and Kulkarni, A. 2024 · 2024
Cited alongside, same era.
GLoRe: when, where, and how to improve LLM reasoning via global and local refinements
Havrilla, A.; Raparthy, S.; Nalmpantis, C.; Dwivedi-Yu, J.; Zhuravynski, M.; Hambro, E.; and Raileanu, R. 2024 · 2024
Cited alongside, same era.
Tulu 3: Pushing frontiers in open language model post-training
Lambert, N.; Morrison, J.; Pyatkin, V.; Huang, S.; Ivison, H.; Brahman, F.; Miranda, L. J. V.; Liu, A.; Dziri, N.; Lyu, S.; et al. 2024 · 2024
Cited alongside, same era.
Chain of code: reasoning with a language model-augmented code emulator
Li, C.; Liang, J.; Zeng, A.; Chen, X.; Hausman, K.; Sadigh, D.; Levine, S.; Fei-Fei, L.; Xia, F.; and Ichter, B. 2024 · 2024
Cited alongside, same era.
SciAgent: Tool-augmented Language Models for Scientific Reasoning
Ma, Y.; Gou, Z.; Hao, J.; Xu, R.; Wang, S.; Pan, L.; Yang, Y.; Cao, Y.; and Sun, A. 2024 · 2024
Cited alongside, same era.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y.; Xia, M.; and Chen, D. 2024 · 2024
Cited alongside, same era.
Large language models: A survey
Minaee, S.; Mikolov, T.; Nikzad, N.; Chenaghlu, M.; Socher, R.; Amatriain, X.; and Gao, J. 2024 · 2024
Cited alongside, same era.
Huang, Z.; Yuan, X.; Ju, Y.; Zhao, J.; and Liu, K. 2025 · 2025
Closest in time.
Flashrag: A modular toolkit for efficient retrieval-augmented generation research
Jin, J.; Zhu, Y.; Dou, Z.; Dong, G.; Yang, X.; Zhang, C.; Zhao, T.; Yang, Z.; and Wen, J.-R. 2025b · 2025
Closest in time.
VinePPO: Refining Credit Assignment in RL Training of LLMs
Kazemnejad, A.; Aghajohari, M.; Portelance, E.; Sordoni, A.; Reddy, S.; Courville, A.; and Le Roux, N. 2025 · 2025
Closest in time.
When thinking fails: The pitfalls of reasoning for instruction-following in llms
Li, X.; Yu, Z.; Zhang, Z.; Chen, X.; Zhang, Z.; Zhuang, Y.; Sadagopan, N.; and Beniwal, A. 2025 · 2025
Closest in time.
Torl: Scaling tool-integrated rl
Li, X.; Zou, H.; and Liu, P. 2025 · 2025
Closest in time.
ToolACE: Winning the Points of LLM Function Calling
Liu, W.; Huang, X.; Zeng, X.; xinlong hao; Yu, S.; Li, D.; Wang, S.; Gan, W.; Liu, Z.; Yu, Y.; WANG, Z.; Wang, Y.; Ning, W.; Hou, Y.; Wang, B.; Wu, C.; Xinzhi, W.; Liu, Y.; Wang, Y.; Tang, D.; Tu, D.; Shang, L.; Jiang, X.; Tang, R.; Lian, D.; Liu, Q.; and Chen, E. 2025 · 2025
Closest in time.
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
Lu, P.; Chen, B.; Liu, S.; Thapa, R.; Boen, J.; and Zou, J. 2025 · 2025
Closest in time.
Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving
Mai, X.; Xu, H.; Wang, W.; Zhang, Y.; Zhang, W.; et al. 2025 · 2025
Closest in time.
Tool learning with large language models: a survey
Qu, C.; Dai, S.; Wei, X.; Cai, H.; Wang, S.; Yin, D.; Xu, J.; and Wen, J.-r. 2025 · 2025
Closest in time.
Hybridflow: A flexible and efficient rlhf framework
Sheng, G.; Zhang, C.; Ye, Z.; Wu, X.; Zhang, W.; Zhang, R.; Peng, Y.; Lin, H.; and Wu, C. 2025 · 2025
Closest in time.
R1-searcher: Incentivizing the search capability in llms via reinforcement learning
Song, H.; Jiang, J.; Min, Y.; Chen, J.; Chen, Z.; Zhao, W. X.; Fang, L.; and Wen, J.-R. 2025 · 2025
Closest in time.
Otc: Optimal tool calls via reinforcement learning
Wang, H.; Qian, C.; Zhong, W.; Chen, X.; Qiu, J.; Huang, S.; Jin, B.; Wang, M.; Wong, K.-F.; and Ji, H. 2025 · 2025
Closest in time.
Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs
Wei, Y.; Yu, X.; Pan, T.; Li, A.; and Du, L. 2025 · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.; Dai, W.; Fan, T.; Liu, G.; Liu, L.; et al. 2025 · 2025
Closest in time.
Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
Zeng, W.; Huang, Y.; Liu, Q.; Liu, W.; He, K.; Ma, Z.; and He, J. 2025 · 2025
Closest in time.
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Zhou, Z.; Qu, A.; Wu, Z.; Kim, S.; Prakash, A.; Rus, D.; Zhao, J.; Low, B. K. H.; and Liang, P. P. 2025 · 2025
Closest in time.