Fetching the paper…
Reading the bibliography…
In the classical Reinforcement Learning from Human Feedback (RLHF) framework, Proximal Policy Optimization (PPO) is employed to learn from sparse, sentence-level rewards -- a challenging scenario in traditional deep reinforcement learning.
On the weaknesses of reinforcement learning for neural machine translation
Choshen, L · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A · 1952
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S · 2002
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Völske, M · 2017
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Earlier work this paper cites.
Improved optimistic algorithms for logistic bandits
Faury, L · 2020
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A · 2021
Earlier work this paper cites.
Preference-based online learning with dueling bandits: A survey
Bengs, V · 2021
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L · 2021
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Jin, Y · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R · 2021
Earlier work this paper cites.
Dueling rl: reinforcement learning with trajectory preferences
Pacchiano, A · 2021
Earlier work this paper cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P · 2021
Earlier work this paper cites.
Optimal algorithms for stochastic contextual preference bandits
Saha, A · 2021
Earlier work this paper cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y · 2022
Earlier work this paper cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S · 2022
Cited alongside, same era.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X · 2022
Cited alongside, same era.
When is partially observable reinforcement learning not scary?
Liu, Q · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
Solving math word problems with process-and outcome-based feedback
Uesato, J · 2022
Cited alongside, same era.
Nearly optimal policy optimization with stable at any time guarantee
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Ahmadian, A · 2024
Closest in time.
Dense reward for free in reinforcement learning from human feedback
Chan, A. J · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
Dong, H · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, T · 2022
Cited alongside, same era.
Xiong, W · 2022
Cited alongside, same era.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Zhong, H · 2022
Cited alongside, same era.
Introducing claude. https://www.anthropic.com/index/introducing-claude
Anthropic · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S · 2023
Cited alongside, same era.
Dubey, A · 2024
Closest in time.
Snorkel-mistral-pairrm-dpo. https://huggingface.co/snorkelai/Snorkel-Mistral-PairRM-DPO
Hoang Tran, B. H., Chris Glaze · 2024
Closest in time.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J · 2024
Closest in time.
On the statistical efficiency of mean-field reinforcement learning with general function approximation
Huang, J · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N · 2024
Closest in time.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Li, T · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y · 2024
Closest in time.
Disentangling length from quality in direct preference optimization
Park, R · 2024
Closest in time.
From r r to q ∗ q^{*} : Your language model is secretly a q-function
Rafailov, R · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y · 2024
Closest in time.
Wang, H · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Wu, Z · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C · 2024
Closest in time.
Token-level direct preference optimization
Zeng, Y · 2024
Closest in time.
A theoretical analysis of optimistic proximal policy optimization in linear markov decision processes
Zhong, H · 2024
Closest in time.
Process reinforcement through implicit rewards
Cui, G · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Hu, J · 2025
Closest in time.
Segmenting text and learning their rewards for improved rlhf in language model
Yin, Y · 2025
Closest in time.