Fetching the paper…
Reading the bibliography…
Offline preference optimization allows fine-tuning large models directly from offline data, and has proved effective in recent alignment practices.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Bernhard E. Boser, Isabelle M. Guyon, and Vladimir N. Vapnik · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
A desicion-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1995
Earlier work this paper cites.
Are loss functions all the same?
Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, and Alessandro Verri · 2004
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L. Bartlett, Michael I. Jordan, and Jon D. McAuliffe · 2006
Earlier work this paper cites.
On the design of loss functions for classification: Yheory, robustness to outliers, and SavageBoost
Hamed Masnadi-Shirazi and Nuno Vasconcelos · 2008
Earlier work this paper cites.
The elements of statistical learning: Data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome H. Friedman · 2009
Earlier work this paper cites.
A view of margin losses as regularizers of probability estimates
Hamed Masnadi-Shirazi and Nuno Vasconcelos · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
VarGrad: A low-variance gradient estimator for variational inference
Lorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco Ruiz, and Omer Deniz Akyildiz · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano · 2020
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Scaling up models and data with t5x and seqio
Adam Roberts, Hyung Won Chung, Gaurav Mishra, Anselm Levskaya, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, et al · 2023
Later among the works it cites.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, et al · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
SLiC-HF: Sequence likelihood calibration with human feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos · 2024
Closest in time.
Human alignment of large language models through online preference optimisation
Daniele Calandriello, Daniel Guo, Remi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, and Bilal Piot · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Nash learning from human feedback
Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, and Bilal Piot · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Closest in time.