Fetching the paper…
Reading the bibliography…
Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting.
Linear lexicographic optimization
H. Isermann · 1982
Earlier work this paper cites.
Decisions with Multiple Objectives: Preferences and Value Tradeoffs
Ralph Keeney and Howard Raiffa · 1993
Earlier work this paper cites.
MOGA: Multi-objective genetic algorithms
Tadahiko Murata and Hisao Ishibuchi · 1995
Earlier work this paper cites.
Nonlinear Multiobjective Optimization
Kaisa Miettinen · 1998
Earlier work this paper cites.
Survey of multi-objective optimization methods for engineering
Timothy Marler and Jasbir Arora · 2004
Earlier work this paper cites.
An EMO algorithm using the hypervolume measure as selection criterion
Michael Emmerich, Nicola Beume, and Boris Naujoks · 2005
Earlier work this paper cites.
Introduction to multiobjective optimization: Interactive approaches
Kaisa Miettinen, Francisco Ruiz, and Wierzbicki · 2008
Earlier work this paper cites.
Multi-Objective Evolutionary Optimisation for Product Design and Manufacturing
Lihui Wang, Amos Ng, and Kalyanmoy Deb · 2011
Earlier work this paper cites.
Bounding the effectiveness of hypervolume-based ( μ \mu + λ \lambda )-archiving algorithms
Tamara Ulrich and Lothar Thiele · 2012
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stephane Boucheron, Gabor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
A survey on multiobjective evolutionary algorithms for the solution of the portfolio optimization problem and other finance and economics applications
Antonin Ponsich, Antonio Jaimes, and Carlos Coello · 2013
Earlier work this paper cites.
A multi-objective optimization model for sustainable logistics facility location
Tang Xifeng, Zhang Ji, and Xu Peng · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Pareto front identification from stochastic bandit feedback
Peter Auer, Chao-Kai Chiang, Ronald Ortner, and Madalina Drugan · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
A tutorial on multiobjective optimization: Fundamentals and evolutionary methods
Michael Emmerich and Andre Deutz · 2018
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Differentiable expected hypervolume improvement for parallel multi-objective bayesian optimization
Samuel Daulton, Maximilian Balandat, and Eytan Bakshy · 2020
Cited alongside, same era.
Deep reinforcement learning for multiobjective optimization
Kaiwen Li, Tao Zhang, and Rui Wang · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Random hypervolume scalarizations for provable multi-objective black box optimization
Richard Zhang and Daniel Golovin · 2020
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao · 2023
Later among the works it cites.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Jing Shao, Xiangyu Yue, Chao Yang, Wanli Ouyang, and Yu Qiao · 2023
Later among the works it cites.
Deal: Decoding-time alignment for large language models
James Y Huang, Sailik Sengupta, Daniele Bonadiman, Yi-an Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchoff, and Dan Roth · 2024
Closest in time.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2024
Closest in time.
Rui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu, Han Zhong, Dong Yu, and Jianshu Chen · 2024
Closest in time.