Fetching the paper…
Reading the bibliography…
In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs).
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Semi-supervised classification by low density separation
Olivier Chapelle and Alexander Zien · 2005
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Truncated importance sampling
Edward L Ionides · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee et al · 2013
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Earlier work this paper cites.
Multi-agent systems: A survey
Ali Dorri, Salil S Kanhere, and Raja Jurdak · 2018
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
Xin Wang, Wenhan Xiong, Hongmin Wang, and William Yang Wang · 2018
Earlier work this paper cites.
Mean field multi-agent reinforcement learning
Yaodong Yang, Rui Luo, Minne Li, Ming Zhou, Weinan Zhang, and Jun Wang · 2018
Earlier work this paper cites.
A meta-mdp approach to exploration for lifelong reinforcement learning
Francisco Garcia and Philip S Thomas · 2019
Earlier work this paper cites.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond J Mooney · 2019
Earlier work this paper cites.
Learning agile and dynamic motor skills for legged robots
Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, and Marco Hutter · 2019
Earlier work this paper cites.
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Learning to utilize shaping rewards: A new approach of reward shaping
Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang, Yingfeng Chen, Jianye Hao, Feng Wu, and Changjie Fan · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder De Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec · 2020
Earlier work this paper cites.
Comps: Continual meta policy search
Glen Berseth, Zhiwei Zhang, Grace Zhang, Chelsea Finn, and Sergey Levine · 2021
Earlier work this paper cites.
Multi-agent reinforcement learning: A review of challenges and applications
Lorenzo Canese, Gian Carlo Cardarilli, Luca Di Nunzio, Rocco Fazzolari, Daniele Giardino, Marco Re, and Sergio Spanò · 2021
Earlier work this paper cites.
Maximum entropy rl (provably) solves some robust rl problems
Benjamin Eysenbach and Sergey Levine · 2021
Earlier work this paper cites.
Dynamic neural networks: A survey
Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang · 2021
Earlier work this paper cites.
Temporal-logic-based reward shaping for continuing reinforcement learning tasks
Yuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah, Ufuk Topcu, and Peter Stone · 2021
Earlier work this paper cites.
Sler: Self-generated long-term experience replay for continual reinforcement learning
Chunmao Li, Yang Li, Yinliang Zhao, Peng Peng, and Xupeng Geng · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S Sutton · 2021
Earlier work this paper cites.
Continual world: A robotic benchmark for continual reinforcement learning
Maciej Wołczyk, Michał Zając, Razvan Pascanu, Łukasz Kuciński, and Piotr Miłoś · 2021
Earlier work this paper cites.
Reincarnating reinforcement learning: Reusing prior computation to accelerate progress
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2022
Earlier work this paper cites.
Building a subspace of policies for scalable continual learning
Jean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier, Ludovic Denoyer, and Roberta Raileanu · 2022
Earlier work this paper cites.
Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity
Abhishek Gupta, Aldo Pacchiano, Yuexiang Zhai, Sham Kakade, and Sergey Levine · 2022
Earlier work this paper cites.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Earlier work this paper cites.
Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Runze Liu, Fengshuo Bai, Yali Du, and Yaodong Yang · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T Connor, Neil Burch, Thomas Anthony, et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Earlier work this paper cites.
Will we run out of data? limits of llm scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn · 2022
Earlier work this paper cites.
Scienceworld: Is your agent smarter than a 5th grader?
Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Earlier work this paper cites.
Lifelong reinforcement learning with temporal logic formulas and reward machines
Xuejing Zheng, Chao Yu, and Minjie Zhang · 2022
Earlier work this paper cites.
Scaling laws for generative mixed-modal language models
Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer · 2023
Earlier work this paper cites.
Autonomous chemical research with large language models
Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes · 2023
Earlier work this paper cites.
Settling the reward hypothesis
Michael Bowling, John D Martin, David Abel, and Will Dabney · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Earlier work this paper cites.
Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo Perez-Vicente, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
Earlier work this paper cites.
Entity divider with language grounding in multi-agent reinforcement learning
Ziluo Ding, Wanpeng Zhang, Junpeng Yue, Xiangjun Wang, Tiejun Huang, and Zongqing Lu · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Earlier work this paper cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Earlier work this paper cites.
Language models, agent models, and world models: The law for machine reasoning and planning
Zhiting Hu and Tianmin Shu · 2023
Earlier work this paper cites.
Multi-agent reinforcement learning: A comprehensive survey
Dom Huh and Prasant Mohapatra · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Assessing the limits of zero-shot foundation models in single-cell biology, 2023
Kasia Z. Kedzierska, Lorin Crawford, Ava P. Amini, and Alex X. Lu · 2023
Earlier work this paper cites.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Earlier work this paper cites.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, Aske Plaat, Catholijn M Jonker, et al · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Earlier work this paper cites.
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Text2reward: Reward shaping with language models for reinforcement learning
Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Qian Luo, Victor Zhong, Yanchao Yang, and Tao Yu · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al · 2023
Earlier work this paper cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Earlier work this paper cites.
Critique-out-loud reward models
Zachary Ankner, Mansheej Paul, Brandon Cui, Jonathan D Chang, and Prithviraj Ammanabrolu · 2024
Earlier work this paper cites.
Zero-shot model-based reinforcement learning using large language models
Abdelhakim Benechehab, Youssef Attia El Hili, Ambroise Odonnat, Oussama Zekri, Albert Thomas, Giuseppe Paolo, Maurizio Filippone, Ievgen Redko, and Balázs Kégl · 2024
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Earlier work this paper cites.
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Earlier work this paper cites.
How to build the virtual cell with artificial intelligence: Priorities and opportunities
Charlotte Bunne, Yusuf Roohani, Yanay Rosen, Ankit Gupta, Xikun Zhang, Marcel Roed, Theo Alexandrov, Mohammed AlQuraishi, Patricia Brennan, Daniel B Burkhardt, et al · 2024
Earlier work this paper cites.
Mle-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al · 2024
Earlier work this paper cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al · 2024
Earlier work this paper cites.
Loss of plasticity in deep continual learning
Shibhansh Dohare, J Fernando Hernandez-Garcia, Qingfeng Lan, Parash Rahman, A Rupam Mahmood, and Richard S Sutton · 2024
Earlier work this paper cites.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Earlier work this paper cites.
A comprehensive survey of reinforcement learning: From algorithms to practical challenges
Majid Ghasemi, Amir Hossein Moosavi, and Dariush Ebrahimi · 2024
Earlier work this paper cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Earlier work this paper cites.
Is your llm secretly a world model of the internet? model-based planning for web agents
Yu Gu, Kai Zhang, Yuting Ning, Boyuan Zheng, Boyu Gou, Tianci Xue, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, et al · 2024
Earlier work this paper cites.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Earlier work this paper cites.
Intuitive fine-tuning: Towards simplifying alignment into a single process
Ermo Hua, Biqing Qi, Kaiyan Zhang, Yue Yu, Ning Ding, Xingtai Lv, Kai Tian, and Bowen Zhou · 2024
Earlier work this paper cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Earlier work this paper cites.
Towards efficient exact optimization of language model alignment
Haozhe Ji, Cheng Lu, Yilin Niu, Pei Ke, Hongning Wang, Jun Zhu, Jie Tang, and Minlie Huang · 2024
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al · 2024
Cited alongside, same era.
Tulu 3: Pushing frontiers in open language model post-training
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al · 2024
Cited alongside, same era.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2024
Cited alongside, same era.
Fine-tuning vision-language-action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang · 2025
Closest in time.
Andrew Kiruluta, Andreas Lemos, and Priscilla Burity · 2025
Closest in time.
Mercury: Ultra-fast language models based on diffusion
Inception Labs, Samar Khanna, Siddhant Kharbanda, Shufan Li, Harshit Varma, Eric Wang, Sawyer Birnbaum, Ziyang Luo, Yanis Miraoui, Akash Palrecha, et al · 2025
Closest in time.
Computerrl: Scaling end-to-end online reinforcement learning for computer use agents, 2025
Hanyu Lai, Xiao Liu, Yanxiao Zhao, Han Xu, Hanchen Zhang, Bohao Jing, Yanyu Ren, Shuntian Yao, Yuxiao Dong, and Jie Tang · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Cited alongside, same era.
The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning
Suyun Liu and Luis Nunes Vicente · 2024
Cited alongside, same era.
Autopsv: Automated process-supervised verifier
Jianqiao Lu, Zhiyang Dou, Hongru Wang, Zeyu Cao, Jianbo Dai, Yunlong Feng, and Zhijiang Guo · 2024
Cited alongside, same era.
A survey on model-based reinforcement learning
Fan-Ming Luo, Tian Xu, Hang Lai, Xiong-Hui Chen, Weinan Zhang, and Yang Yu · 2024
Cited alongside, same era.
Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fränken, Chelsea Finn, and Alon Albalak · 2024
Cited alongside, same era.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, et al · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al · 2024
Cited alongside, same era.
Training software engineering agents and verifiers with swe-gym
Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang · 2024
Cited alongside, same era.
Jiacheng Lin and Zhenbang Wu · 2025
Closest in time.
Jun Ling, Yao Qi, Tao Huang, Shibo Zhou, Yanqin Huang, Jiang Yang, Ziqi Song, Ying Zhou, Yang Yang, Heng Tao Shen, et al · 2025
Closest in time.
Code-r1: Reproducing r1 for code with reliable rewards
Jiawei Liu and Lingming Zhang · 2025
Closest in time.
Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models
Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, and Yi Dong · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang · 2025
Closest in time.
Cpgd: Toward stable rule-based reinforcement learning for language models
Zongkai Liu, Fanqing Meng, Lingxiao Du, Zhixiang Zhou, Chao Yu, Wenqi Shao, and Qiaosheng Zhang · 2025
Closest in time.
Adsqa: Towards advertisement video understanding
Xinwei Long, Kai Tian, Peng Xu, Guoli Jia, Jingxuan Li, Sa Yang, Yihua Shao, Kaiyan Zhang, Che Jiang, Hao Xu, et al · 2025
Closest in time.
Towards a unified view of large language model post-training
Xingtai Lv, Yuxin Zuo, Youbang Sun, Hongyi Liu, Yuntian Wei, Zhekai Chen, Lixuan He, Xuekai Zhu, Kaiyan Zhang, Bingning Wang, et al · 2025
Closest in time.
Exploring the limit of outcome reward for learning mathematical reasoning
Chengqi Lyu, Songyang Gao, Yuzhe Gu, Wenwei Zhang, Jianfei Gao, Kuikun Liu, Ziyi Wang, Shuaibin Li, Qian Zhao, Haian Huang, et al · 2025
Closest in time.
Rl squeezes, sft expands: A comparative study of reasoning llms
Kohsei Matsutani, Shota Takashiro, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, and Yutaka Matsuo · 2025
Closest in time.
Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1, 2025
Justus Mattern, Sami Jaghouar, Manveer Basra, Jannik Straube, Matthew Di Ferrante, Felix Gabriel, Jack Min Ong, Vincent Weisser, and Johannes Hagemann · 2025
Closest in time.
O2-searcher: A searching-based agent model for open-domain open-ended question answering
Jianbiao Mei, Tao Hu, Daocheng Fu, Licheng Wen, Xuemeng Yang, Rong Wu, Pinlong Cai, Xinyu Cai, Xing Gao, Yu Yang, et al · 2025
Closest in time.
Ivan Moshkov, Darragh Hanley, Ivan Sorokin, Shubham Toshniwal, Christof Henkel, Benedikt Schifferer, Wei Du, and Igor Gitman · 2025
Closest in time.
Reinforcement learning finetunes small subnetworks in large language models
Sagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, and Hao Peng · 2025
Closest in time.
Mle-star: Machine learning engineering agent via search and targeted refinement
Jaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin, Sercan Ö Arık, and Tomas Pfister · 2025
Closest in time.
Training a scientific reasoning model for chemistry
Siddharth M. Narayanan, James D. Braza, Ryan-Rhys Griffiths, Albert Bou, Geemi Wellawatte, Mayk Caldas Ramos, Ludovico Mitchener, Samuel G. Rodriques, and Andrew D. White · 2025
Closest in time.
Mlgym: A new framework and benchmark for advancing ai research agents
Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav Vorotilov, Gaurav Chaurasia, et al · 2025
Closest in time.
Sfr-deepresearch: Towards effective reinforcement learning for autonomously reasoning single agents
Xuan-Phi Nguyen, Shrey Pandit, Revanth Gangi Reddy, Austin Xu, Silvio Savarese, Caiming Xiong, and Shafiq Joty · 2025
Closest in time.
Large language diffusion models, 2025
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li · 2025
Closest in time.
Predictive scaling laws for efficient grpo training of large reasoning models
Datta Nimmaturi, Vaishnavi Bhargava, Rajat Ghosh, Johnu George, and Debojyoti Dutta · 2025
Closest in time.
Virtual cells: Predict, explain, discover, 2025
Emmanuel Noutahi, Jason Hartford, Prudencio Tossou, Shawn Whitfield, Alisandra K. Denton, Cas Wognum, Kristina Ulicna, Michael Craig, Jonathan Hsu, Michael Cuccarese, et al · 2025
Closest in time.
Alphaevolve: A coding agent for scientific and algorithmic discovery
Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al · 2025
Closest in time.
Humza Nusrat · 2025
Closest in time.
Nemo rl: A scalable and efficient post-training library
NVIDIA-NeMo · 2025
Closest in time.
Laviplan: Language-guided visual path planning with rlvr
Hayeon Oh · 2025
Closest in time.
Introducing gpt-4o image generation
OpenAI · 2025
Closest in time.
Spacer: Reinforcing mllms in video spatial reasoning
Kun Ouyang, Yuanxin Liu, Haoning Wu, Yi Liu, Hao Zhou, Jie Zhou, Fandong Meng, and Xu Sun · 2025
Closest in time.
Tinyzero
Jiayi Pan, Junjie Zhang, Xingyao Wang, Lifan Yuan, Hao Peng, and Alane Suhr · 2025
Closest in time.
Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse
Zhenyu Pan and Han Liu · 2025
Closest in time.
Advancing slm tool-use capability using reinforcement learning
Dhruvi Paprunia, Vansh Kharidia, and Pankti Doshi · 2025
Closest in time.
Curriculum reinforcement learning from easy to hard tasks improves llm reasoning
Shubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling, Sushil Vemuri, Blake Olson, Eric Li, Yu Zhang, James Caverlee, Dileep Kalathil, et al · 2025
Closest in time.
Chanwoo Park, Seungju Han, Xingzhi Guo, Asuman Ozdaglar, Kaiqing Zhang, and Joo-Kyung Kim · 2025
Closest in time.
Codeforces cots
Guilherme Penedo, Anton Lozhkov, Hynek Kydlíček, Loubna Ben Allal, Edward Beeching, Agustín Piqueres Lajarín, Quentin Gallouédec, Nathan Habib, Lewis Tunstall, and Leandro von Werra · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, and et al · 2025
Closest in time.
Mohammadreza Pourreza, Shayan Talaei, Ruoxi Sun, Xingchen Wan, Hailong Li, Azalia Mirhoseini, Amin Saberi, Sercan Arik, et al · 2025
Closest in time.
Maximizing confidence alone improves reasoning
Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak · 2025
Closest in time.
Synthetic-2 release: Four million collaboratively generated reasoning traces
PrimeIntellect · 2025
Closest in time.
Toolrl: Reward is all tool learning needs
Cheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan Tur, and Heng Ji · 2025
Closest in time.
Mle-dojo: Interactive environments for empowering llm agents in machine learning engineering
Rushi Qiang, Yuchen Zhuang, Yinghao Li, Rongzhi Zhang, Changhao Li, Ian Shu-Hei Wong, Sherry Yang, Percy Liang, Chao Zhang, Bo Dai, et al · 2025
Closest in time.
Supervised fine tuning on curated data is reinforcement learning (and can be improved)
Chongli Qin and Jost Tobias Springenberg · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al · 2025
Closest in time.
Open-medical-r1: How to choose data for rlvr training at medicine domain, 2025
Zhongxi Qiu, Zhang Zhang, Yan Hu, Heng Li, and Jiang Liu · 2025
Closest in time.
Opentable-r1: A reinforcement learning augmented tool agent for open-domain table question answering
Zipeng Qiu · 2025
Closest in time.
Qvq: To see the world with wisdom, 2025
Qwen Team · 2025
Closest in time.
Magistral
Abhinav Rastogi, Albert Q. Jiang, Andy Lo, Gabrielle Berrada, Guillaume Lample, Jason Rute, Joep Barmentlo, Karmesh Yadav, Kartik Khandelwal, Khyathi Raghavi Chandu, et al · 2025
Closest in time.
ZZ Ren, Zhihong Shao, Junxiao Song, Huajian Xin, Haocheng Wang, Wanjia Zhao, Liyue Zhang, Zhe Fu, Qihao Zhu, Dejian Yang, et al · 2025
Closest in time.
Scaling large language models for next-generation single-cell analysis
Syed Asad Rizvi, Daniel Levine, Aakash Patel, Shiyang Zhang, Eric Wang, Sizhuang He, David Zhang, Cerise Tang, Zhuoyang Lyu, Rayyan Darji, Chang Li, Emily Sun, David Jeong, Lawrence Zhao, Jennifer Kwan, David Braun, Brian Hafler, Jeffrey Ishizuka, Rahul M Dhodapkar, Hattie Chung, Shekoofeh Azizi, Bryan Perozzi, and David van Dijk · 2025
Closest in time.
Tapered off-policy reinforce: Stable and efficient reinforcement learning for llms
Nicolas Le Roux, Marc G Bellemare, Jonathan Lebensold, Arnaud Bergeron, Joshua Greaves, Alex Fréchette, Carolyne Pelletier, Eric Thibodeau-Laufer, Sándor Toth, and Sam Work · 2025
Closest in time.
Gaia-2: A controllable multi-view generative world model for autonomous driving
Lloyd Russell, Anthony Hu, Lorenzo Bertoni, George Fedoseev, Jamie Shotton, Elahe Arani, and Gianluca Corrado · 2025
Closest in time.
Vision-language-action models: Concepts, progress, applications and challenges
Ranjan Sapkota, Yang Cao, Konstantinos I Roumeliotis, and Manoj Karkee · 2025
Closest in time.
Llms are greedy agents: Effects of rl fine-tuning on decision-making abilities
Thomas Schmied, Jörg Bornschein, Jordi Grau-Moya, Markus Wulfmeier, and Razvan Pascanu · 2025
Closest in time.
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, et al · 2025
Closest in time.
e3: Learning to explore enables extrapolation of test-time compute for llms
Amrith Setlur, Matthew YR Yang, Charlie Snell, Jeremy Greer, Ian Wu, Virginia Smith, Max Simchowitz, and Aviral Kumar · 2025
Closest in time.
Sem: Reinforcement learning for search-efficient large language models
Zeyang Sha, Shiwen Cui, and Weiqiang Wang · 2025
Closest in time.
Can large reasoning models self-train?
Sheikh Shafayat, Fahim Tajwar, Ruslan Salakhutdinov, Jeff Schneider, and Andrea Zanette · 2025
Closest in time.
Stepfun-prover preview: Let’s think and verify step by step
Shijie Shang, Ruosi Wan, Yue Peng, Yutong Wu, Xiong-hui Chen, Jie Yan, and Xiangyu Zhang · 2025
Closest in time.
Spurious rewards: Rethinking training signals in rlvr
Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, et al · 2025
Closest in time.
Rl’s razor: Why online reinforcement learning forgets less
Idan Shenfeld, Jyothish Pari, and Pulkit Agrawal · 2025
Closest in time.
Hybridflow: A flexible and efficient rlhf framework
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu · 2025
Closest in time.
Sample more to think less: Group filtered policy optimization for concise reasoning
Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, and Dimitris Papailiopoulos · 2025
Closest in time.
Welcome to the era of experience
David Silver and Richard S Sutton · 2025
Closest in time.
A general framework for inference-time scaling and steering of diffusion models
Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath · 2025
Closest in time.
The illusion of diminishing returns: Measuring long horizon execution in llms
Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping · 2025
Closest in time.
A technical survey of reinforcement learning techniques for large language models
Saksham Sahai Srivastava and Vaneet Aggarwal · 2025
Closest in time.
Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards
Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kaddour, and Andreas Köpf · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, et al · 2025
Closest in time.
A large language model-driven reward design framework via dynamic feedback for reinforcement learning
Shengjie Sun, Runze Liu, Jiafei Lyu, Jing-Wen Yang, Liangpeng Zhang, and Xiu Li · 2025
Closest in time.
All roads lead to likelihood: The value of reinforcement learning in fine-tuning
Gokul Swamy, Sanjiban Choudhury, Wen Sun, Zhiwei Steven Wu, and J Andrew Bagnell · 2025
Closest in time.
Tess 2: A large-scale generalist diffusion language model
Jaesung Tae, Hamish Ivison, Sachin Kumar, and Arman Cohan · 2025
Closest in time.
Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy
Hongze Tan and Jianfei Pan · 2025
Closest in time.
Agent kb: Leveraging cross-domain experience for agentic problem solving
Xiangru Tang, Tianrui Qin, Tianhao Peng, Ziyang Zhou, Daniel Shao, Tingting Du, Xinming Wei, Peng Xia, Fang Wu, He Zhu, et al · 2025
Closest in time.
Webshaper: Agentically data synthesizing via information-seeking formalization
Zhengwei Tao, Jialong Wu, Wenbiao Yin, Junkai Zhang, Baixuan Li, Haiyang Shen, Kuan Li, Liwen Zhang, Xinyu Wang, Yong Jiang, et al · 2025
Closest in time.
slime: An sglang-native post-training framework for rl scaling
THUDM · 2025
Closest in time.
Ego-r1: Chain-of-tool-thought for ultra-long egocentric video reasoning
Shulin Tian, Ruiqi Wang, Hongming Guo, Penghao Wu, Yuhao Dong, Xiuying Wang, Jingkang Yang, Hao Zhang, Hongyuan Zhu, and Ziwei Liu · 2025
Closest in time.
avatarl: training language models from scratch with pure reinforcement learning, 2025
tokenbender · 2025
Closest in time.
Post-training large language models via reinforcement learning from self-feedback
Carel van Niekerk, Renato Vukovic, Benjamin Matthias Ruppik, Hsien-chin Lin, and Milica Gašić · 2025
Closest in time.
How to train your llm web agent: A statistical diagnosis
Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Megh Thakkar, Thibault Le Sellier de Chezelles, Nicolas Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, et al · 2025
Closest in time.
Checklists are better than reward models for aligning language models
Vijay Viswanathan, Yanchao Sun, Shuang Ma, Xiang Kong, Meng Cao, Graham Neubig, and Tongshuang Wu · 2025
Closest in time.
Pass@k policy optimization: Solving harder reinforcement learning problems
Christian Walder and Deep Karkhanis · 2025
Closest in time.
Rema: Learning to meta-think for llms with multi-agent reinforcement learning
Ziyu Wan, Yunxiang Li, Xiaoyu Wen, Yan Song, Hanjing Wang, Linyi Yang, Mark Schmidt, Jun Wang, Weinan Zhang, Shuyue Hu, et al · 2025
Closest in time.
Kimina-prover preview: Towards large formal reasoning models with reinforcement learning, 2025
Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys, Junqi Liu, MD Santos, Flood Sung, Marina Vinyes, Zhenzhe Ying, Zekai Zhu, et al · 2025
Closest in time.
Hanyin Wang · 2025
Closest in time.
Look before you leap: A gui-critic-r1 model for pre-operative error diagnosis in gui automation
Yuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Jiabo Ye, Yutong Kou, Ming Yan, Fei Huang, Xiaoshan Yang, et al · 2025
Closest in time.
The asymmetry of verification and verifier’s law
Jason Wei · 2025
Closest in time.
J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning
Chenxi Whitehouse, Tianlu Wang, Ping Yu, Xian Li, Jason Weston, Ilia Kulikov, and Swarnadeep Saha · 2025
Closest in time.
Sailing by the stars: A survey on reward models and learning strategies for learning from rewards
Xiaobao Wu · 2025
Closest in time.
Just enough thinking: Efficient reasoning with adaptive length penalties reinforcement learning
Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, and Nick Haber · 2025
Closest in time.
Mimo: Unlocking the reasoning potential of language model–from pretraining to posttraining
LLM Xiaomi, Bingquan Xia, Bowen Shen, Dawei Zhu, Di Zhang, Gang Wang, Hailin Zhang, Huaqiu Liu, Jiebao Xiao, Jinhao Dong, et al · 2025
Closest in time.
Rihui Xin, Han Liu, Zecheng Wang, Yupeng Zhang, Dianbo Sui, Xiaolin Hu, and Bingning Wang · 2025
Closest in time.
Caprl: Stimulating dense image caption capabilities via reinforcement learning
Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, and Dahua Lin · 2025
Closest in time.
Huihui Xu and Yuanpeng Nie · 2025
Closest in time.
Single-stream policy optimization, 2025
Zhongwen Xu and Zihan Ding · 2025
Closest in time.
Ar2: Adversarial reinforcement learning for abstract reasoning in large language models
Cheng-Kai Yeh, Hsing-Wang Lee, Chung-Hung Kuo, and Hen-Hsen Huang · 2025
Closest in time.
Dynamic and generalizable process reward modeling
Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng, Xipeng Qiu, and Xuanjing Huang · 2025
Closest in time.
ByoungJun Jeon Yooseok Lim · 2025
Closest in time.
Rl tango: Reinforcing generator and verifier together for language reasoning
Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong, Duane S Boning, and Dina Katabi · 2025
Closest in time.
Jixiao Zhang and Chunsheng Zuo · 2025
Closest in time.
Medgr²: Breaking the data barrier for medical reasoning via generative reward learning
Weihai Zhi, Jiayan Guo, and Shangyang Li · 2025
Closest in time.
Glm-4.6: Advanced agentic, reasoning and coding capabilities, 2025
Zhipu-AI · 2025
Closest in time.
A survey on vision-language-action models: An action tokenization perspective
Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, et al · 2025
Closest in time.
Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms
Jiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu, Ke Shen, Jingrui He, and Mengdi Wang · 2025
Closest in time.
Adam Zweiger, Jyothish Pari, Han Guo, Ekin Akyürek, Yoon Kim, and Pulkit Agrawal · 2025
Closest in time.