Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inherited from pre-training.
Asymptotic evaluation of certain markov process expectations for large time. iv
Monroe D Donsker and SR Srinivasa Varadhan · 1983
Earlier work this paper cites.
Practical issues in temporal difference learning
Gerald Tesauro · 1991
Earlier work this paper cites.
A game of prediction with expert advice
Vladimir G Vovk · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov, Frank Hutter, et al · 2017
Earlier work this paper cites.
Probabilistic planning with sequential monte carlo methods
Alexandre Piché, Valentin Thomas, Cyril Ibrahim, Yoshua Bengio, and Chris Pal · 2018
Earlier work this paper cites.
Deep reinforcement learning and the deadly triad
Hado Van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Earlier work this paper cites.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc Bellemare, and Rémi Munos · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Earlier work this paper cites.
A comparative analysis of expected and distributional reinforcement learning
Clare Lyle, Marc G Bellemare, and Pablo Samuel Castro · 2019
Earlier work this paper cites.
A modern introduction to online learning
Francesco Orabona · 2019
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Earlier work this paper cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Quality: Question answering with long input texts, yes!
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, et al · 2021
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Value augmented sampling for language model alignment and personalization
Seungwook Han, Idan Shenfeld, Akash Srivastava, Yoon Kim, and Pulkit Agrawal · 2024
Later among the works it cites.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Dylan J Foster and Akshay Krishnamurthy · 2021
Cited alongside, same era.
An exponential lower bound for linearly realizable mdp with constant suboptimality gap
Yuanhao Wang, Ruosong Wang, and Sham Kakade · 2021
Cited alongside, same era.
Offline reinforcement learning: Fundamental barriers for value function approximation
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Cited alongside, same era.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Cited alongside, same era.
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Cited alongside, same era.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Cited alongside, same era.
Carles Domingo-Enrich, Michal Drozdzal, Brian Karrer, and Ricky TQ Chen · 2024
Later among the works it cites.
Derivative-free guidance in continuous and discrete diffusion models with soft value-based decoding
Xiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia, Gokcen Eraslan, Surag Nair, Tommaso Biancalani, Shuiwang Ji, Aviv Regev, Sergey Levine, et al · 2024
Later among the works it cites.
The central role of the loss function in reinforcement learning
Kaiwen Wang, Nathan Kallus, and Wen Sun · 2024
Later among the works it cites.
More benefits of being distributional: Second-order bounds for reinforcement learning
Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus, and Wen Sun · 2024
Later among the works it cites.
Entropy-regularized process reward model
Hanning Zhang, Pengcheng Wang, Shizhe Diao, Yong Lin, Rui Pan, Hanze Dong, Dylan Zhang, Pavlo Molchanov, and Tong Zhang · 2024
Later among the works it cites.
Learning to achieve goals with belief state transformers
Edward S Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, and John Langford · 2024
Later among the works it cites.
Matching the statistical query lower bound for k k -sparse parity problems with sign stochastic gradient descent
Yiwen Kou, Zixiang Chen, Quanquan Gu, and Sham Kakade · 2024
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Later among the works it cites.
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho, He He, Sainbayar Sukhbaatar, and Jason Weston · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Switching the loss reduces the cost in batch reinforcement learning
Alex Ayoub, Kaiwen Wang, Vincent Liu, Samuel Robertson, James McInerney, Dawen Liang, Nathan Kallus, and Csaba Szepesvari · 2024
Later among the works it cites.
Computationally efficient rl under linear bellman completeness for deterministic dynamics
Runzhe Wu, Ayush Sekhari, Akshay Krishnamurthy, and Wen Sun · 2024
Later among the works it cites.
The power of resets in online reinforcement learning
Zakaria Mhammedi, Dylan J Foster, and Alexander Rakhlin · 2024
Later among the works it cites.
Probabilistic inference in language models via twisted sequential monte carlo
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Grosse · 2024
Later among the works it cites.
Stop regressing: Training value functions via classification for scalable deep rl
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, et al · 2024
Later among the works it cites.
Risk-sensitive rl with optimized certainty equivalents via reduction to standard rl
Kaiwen Wang, Dawen Liang, Nathan Kallus, and Wen Sun · 2024
Later among the works it cites.
Distributional reinforcement learning with regularized wasserstein loss
Ke Sun, Yingnan Zhao, Wulong Liu, Bei Jiang, and Linglong Kong · 2024
Later among the works it cites.
Tuning language models by proxy
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith · 2024
Later among the works it cites.
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason Weston · 2024
Later among the works it cites.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Tengyang Xie, Dylan J Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, and Alexander Rakhlin · 2024
Later among the works it cites.
Vpo: Leveraging the number of votes in preference optimization
Jae Hyeon Cho, Minkyung Park, and Byung-Jun Lee · 2024
Later among the works it cites.
Self-exploring language models: Active preference elicitation for online alignment
Shenao Zhang, Donghan Yu, Hiteshi Sharma, Han Zhong, Zhihan Liu, Ziyi Yang, Shuohang Wang, Hany Hassan, and Zhaoran Wang · 2024
Later among the works it cites.
Edward S Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Jayden Teoh, Bryon Xu, David Yan, Dinesh Jayaraman, et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
Zhaolin Gao, Jonathan Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, Drew Bagnell, Jason D Lee, and Wen Sun · 2025
Closest in time.