Fetching the paper…
Reading the bibliography…
Large language models in the past have typically relied on some form of reinforcement learning with human feedback (RLHF) to better align model responses with human preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
A review of quasi-convex functions
Harvey Greenberg and William Pierskalla · 1971
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Jianqing Fan and Runze Li · 2001
Earlier work this paper cites.
Relative convexity
Jason Palmer · 2003
Earlier work this paper cites.
Subset selection in noise based on diversity measure minimization
B.D. Rao, K. Engan, S. F. Cotter, J. Palmer, and K. Kreutz-Delgado · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
C.M. Bishop · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Enhancing sparsity by reweighted l1 minimization
Emmanuel J Candes, Michael B Wakin, and Stephen P Boyd · 2008
Earlier work this paper cites.
Iteratively reweighted algorithms for compressive sensing
Rick Chartrand and Wotao Yin · 2008
Earlier work this paper cites.
Learning to summarize from human feedback, 2020
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2009
Earlier work this paper cites.
Iterative reweighted ℓ 1 \ell_{1} and ℓ 2 \ell_{2} methods for finding sparse solutions
David Wipf and Srikantan Nagarajan · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Revisiting Bayesian blind deconvolution
David Wipf and Haichao Zhang · 2014
Earlier work this paper cites.
Strong NP-hardness for sparse optimization with concave penalty functions
Yichen Chen, Dongdong Ge, Mengdi Wang, Zizhuo Wang, Yinyu Ye, and Hao Yin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Compressing neural networks using the variational information bottleneck
Bin Dai, Chen Zhu, Baining Guo, and David Wipf · 2018
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Practical and consistent estimation of f-divergences
Paul Rubenstein, Olivier Bousquet, Josip Djolonga, Carlos Riquelme, and Ilya O Tolstikhin · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
Direct preference optimization with an offset
Afra Amini, Tim Vieira, and Ryan Cotterell · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2024
Closest in time.
Towards analyzing and understanding the limitations of dpo: A theoretical perspective
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Cited alongside, same era.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed · 2023
Cited alongside, same era.
Duanyu Feng, Bowen Qin, Chen Huang, Zheng Zhang, and Wenqiang Lei · 2024
Closest in time.
Learn your reference model for real good alignment
Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov, Nikita Surnachev, Yaroslav Aksenov, Ian Maksimov, Nikita Balagansky, and Daniil Gavrilov · 2024
Closest in time.
Orpo: Monolithic preference optimization without reference model
Jiwoo Hong, Noah Lee, and James Thorne · 2024
Closest in time.
Understanding the learning dynamics of alignment with human feedback
Shawn Im and Yixuan Li · 2024
Closest in time.
Policy optimization in rlhf: The impact of out-of-preference data
Ziniu Li, Tian Xu, and Yang Yu · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White · 2024
Closest in time.
Disentangling length from quality in direct preference optimization
Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn, and Aviral Kumar · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Closest in time.
Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints
Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen · 2024
Closest in time.