Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks.
Zur theorie der gesellschaftsspiele
J Von Neumann · 1928
Earlier work this paper cites.
Non-cooperative games
John F Nash et al · 1950
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Infinite Dimensional Analysis: a Hitchhiker’s Guide
Charalambos D. Aliprantis and Kim C. Border · 2006
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Earlier work this paper cites.
Airdialogue: An environment for goal-oriented dialogue research
Wei Wei, Quoc Le, Andrew Dai, and Jia Li · 2018
Earlier work this paper cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Earlier work this paper cites.
A modern introduction to online learning
Francesco Orabona · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
What matters in on-policy reinforcement learning? a large-scale empirical study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphael Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2020
Cited alongside, same era.
On the theory of reinforcement learning with once-per-episode feedback
Niladri Chatterji, Aldo Pacchiano, Peter Bartlett, and Michael Jordan · 2021
Cited alongside, same era.
Nash learning from human feedback
Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, et al · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online markov decision processes with aggregate bandit feedback
Alon Cohen, Haim Kaplan, Tomer Koren, and Yishay Mansour · 2021
Cited alongside, same era.
Reinforcement learning with trajectory feedback
Yonathan Efroni, Nadav Merlis, and Shie Mannor · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, and Liwei Wang · 2022
Cited alongside, same era.
What would jiminy cricket do? towards agents that behave morally
Dan Hendrycks, Christine Zhu, Mantas Mazeika, Jesus Navarro, Dawn Song, Andy Zou, Bo Li, Sahil Patel, and Jacob Steinhardt · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Scienceworld: Is your agent smarter than a 5th grader?
Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu · 2022
Cited alongside, same era.
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Is RLHF more difficult than standard RL? a theoretical perspective
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Later among the works it cites.
Making rl with preference-based feedback efficient via randomization
Runzhe Wu and Wen Sun · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
Human alignment of large language models through online preference optimisation
Daniele Calandriello, Daniel Guo, Remi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, et al · 2024
Closest in time.
Near-optimal regret in linear mdps with aggregate bandit feedback
Asaf Cassel, Haipeng Luo, Aviv Rosenberg, and Dmitry Sotnikov · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Google · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
From r r to Q ⋆ Q^{\star} : Your language model is secretly a q-function
Rafael Rafailov, Joey Hejna, Ryan Park, and Chelsea Finn · 2024
Closest in time.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment, 2024
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Closest in time.