Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Empirical Processes in M-Estimation
Sara A. van de Geer · 2000
Earlier work this paper cites.
From ϵ \epsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
A unified view of entropy-regularized Markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
On oracle-efficient PAC RL with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, and Sham M Kakade · 2019
Earlier work this paper cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Earlier work this paper cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Earlier work this paper cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Earlier work this paper cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Earlier work this paper cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Earlier work this paper cites.
Preference-based reinforcement learning with finite-time guarantees
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh, and Artur Dubrawski · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Bilinear classes: A structural framework for provable generalization in RL
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Earlier work this paper cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Cited alongside, same era.
Dueling RL: reinforcement learning with trajectory preferences
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Value-incentivized preference optimization: A unified approach to online and offline rlhf, 2024
Shicong Cen, Jincheng Mei, Katayoon Goshvadi, Hanjun Dai, Tong Yang, Sherry Yang, Dale Schuurmans, Yuejie Chi, and Bo Dai · 2024
Closest in time.
Dataset reset policy optimization for rlhf
Jonathan D Chang, Wenhao Shan, Owen Oertell, Kianté Brantley, Dipendra Misra, Jason D Lee, and Wen Sun · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Provably sample efficient rlhf via active preference optimization
Nirjhar Das, Souradip Chakraborty, Aldo Pacchiano, and Sayak Ray Chowdhury · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Cited alongside, same era.
Computational-statistical gap in reinforcement learning
Daniel Kane, Sihan Liu, Shachar Lovett, and Gaurav Mahajan · 2022
Cited alongside, same era.
KL-entropy-regularized RL with a generative model is minimax optimal
Tadashi Kozuno, Wenhao Yang, Nino Vieillard, Toshinori Kitamura, Yunhao Tang, Jincheng Mei, Pierre Ménard, Mohammad Gheshlaghi Azar, Michal Valko, Rémi Munos, et al · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
GEC: A unified framework for interactive decision making in MDP, POMDP, and beyond
Han Zhong, Wei Xiong, Sirui Zheng, Liwei Wang, Zhaoran Wang, Zhuoran Yang, and Tong Zhang · 2022
Cited alongside, same era.
Foundations of reinforcement learning and interactive decision making
Dylan J Foster and Alexander Rakhlin · 2023
Cited alongside, same era.
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang · 2024
Closest in time.
Exploration-driven policy optimization in RLHF: Theoretical insights on efficient data utilization
Yihan Du, Anna Winnicki, Gal Dalal, Shie Mannor, and R Srikant · 2024
Closest in time.
Length-controlled AlpacaEval: A simple way to debias automatic evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto · 2024
Closest in time.
Efficient exploration for llms
Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, and Benjamin Van Roy · 2024
Closest in time.
KTO: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
REBEL: Reinforcement learning via regressing relative rewards
Zhaolin Gao, Jonathan D Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J Andrew Bagnell, Jason D Lee, and Wen Sun · 2024
Closest in time.
Noah Golowich, Ankur Moitra, and Dhruv Rohatgi · 2024
Closest in time.
Direct language model alignment from online AI feedback
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao, Bilal Piot, et al · 2024
Closest in time.
From live data to high-quality benchmarks: The Arena-Hard pipeline, April 2024
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Maximize to explore: One objective function fusing estimation, planning, and exploration
Zhihan Liu, Miao Lu, Wei Xiong, Han Zhong, Hao Hu, Shenao Zhang, Sirui Zheng, Zhuoran Yang, and Zhaoran Wang · 2024
Closest in time.
Orca-Math: Unlocking the potential of SLMs in grade school math
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah · 2024
Closest in time.
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho, He He, Sainbayar Sukhbaatar, and Jason Weston · 2024
Closest in time.
From r r to Q ⋆ Q^{\star} : Your language model is secretly a Q-function
Rafael Rafailov, Joey Hejna, Ryan Park, and Chelsea Finn · 2024
Closest in time.
Direct Nash Optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms, 2024
Yunhao Tang, Daniel Zhaohan Guo, Zeyu Zheng, Daniele Calandriello, Yuan Cao, Eugene Tarassov, Rémi Munos, Bernardo Ávila Pires, Michal Valko, Yong Cheng, and Will Dabney · 2024
Closest in time.
snorkelai/snorkel-mistral-pairrm-dpo, 2024
Hoang Tran, Chris Glaze, and Braden Hancock · 2024
Closest in time.
A theoretical analysis of Nash learning from human feedback under general KL-regularized preference
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Closest in time.
Self-exploring language models: Active preference elicitation for online alignment, 2024
Shenao Zhang, Donghan Yu, Hiteshi Sharma, Ziyi Yang, Shuohang Wang, Hany Hassan, and Zhaoran Wang · 2024
Closest in time.