Fetching the paper…
Reading the bibliography…
Direct Preference Optimization (DPO) trains a language model using human preference data, bypassing the explicit reward modeling phase of Reinforcement Learning from Human Feedback (RLHF).
Introduction to Probability, Statistics, and Random Processes
Pishro-Nik, H. 2014 · 2014
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020 · 2020
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Earlier work this paper cites.
Understanding Dataset Difficulty with 𝒱 \mathcal{V} -Usable Information
Ethayarajh, K.; Choi, Y.; and Swayamdipta, S. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q. G.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; et al. 2023 · 2023
Earlier work this paper cites.
UltraFeedback: Boosting Language Models with High-quality Feedback
Cui, G.; Yuan, L.; Ding, N.; Yao, G.; Zhu, W.; Ni, Y.; Xie, G.; Liu, Z.; and Sun, M. 2023 · 2023
Cited alongside, same era.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Dubois, Y.; Li, X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
Human-Aware Loss Functions (HALOs)
Ethayarajh, K.; Xu, W.; Jurafsky, D.; and Kiela, D. 2023 · 2023
Cited alongside, same era.
AlpacaEval: An Automatic Evaluator of Instruction-following Models
Li, X.; Zhang, T.; Dubois, Y.; Taori, R.; Gulrajani, I.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
A note on DPO with noisy preferences & relationship to IPO
Mitchell, E. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources
Wang, Y.; Ivison, H.; Dasigi, P.; Hessel, J.; Khot, T.; Chandu, K.; Wadden, D.; MacMillan, K.; Smith, N. A.; Beltagy, I.; et al. 2023 · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G.; Guo, Z. D.; Piot, B.; Munos, R.; Rowland, M.; Valko, M.; and Calandriello, D. 2024 · 2024
Closest in time.
Provably robust dpo: Aligning language models with noisy feedback
Chowdhury, S. R.; Kini, A.; and Natarajan, N. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
Ethayarajh, K.; Xu, W.; Muennighoff, N.; Jurafsky, D.; and Kiela, D. 2024 · 2024
Closest in time.
ORPO: Monolithic preference optimization without reference model
Hong, J.; Lee, N.; and Thorne, J. 2024 · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A.; Kilcher, Y.; von Rütte, D.; Anagnostidis, S.; Tam, Z. R.; Stevens, K.; Barhoum, A.; Nguyen, D.; Stanley, O.; Nagyfi, R.; et al. 2024 · 2024
Closest in time.