Fetching the paper…
Reading the bibliography…
Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous.
“A law of comparative judgment.”
LL Thurstone · 1927
Earlier work this paper cites.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“Stimulus and response generalization: A stochastic model relating generalization to distance in psychological space”
Roger. Shepard · 1957
Earlier work this paper cites.
“Individual choice behavior”
R Luce · 1959
Earlier work this paper cites.
“The relationship between Luce’s choice axiom, Thurstone’s theory of comparative judgment, and the double exponential distribution”
John Yellott · 1977
Earlier work this paper cites.
“The Nash social welfare function”
Mamoru Kaneko and Kenjiro Nakamura · 1979
Earlier work this paper cites.
“On the epistemic limits of personalized prediction”
Lucas Monteiro, Carol Long, Berk Ustun and Flavio Calmon · 1991
Earlier work this paper cites.
“Boolean functions whose Fourier transform is concentrated on the first two levels”
Ehud Friedgut, Gil Kalai and Assaf Naor · 2002
Earlier work this paper cites.
“Modeling purposeful adaptive behavior with the principle of maximum causal entropy”
Brian Ziebart · 2010
Earlier work this paper cites.
“Truth is a lie: Crowd truth and the seven myths of human annotation”
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
“Inherent disagreements in human textual inferences”
Ellie Pavlick and Tom Kwiatkowski · 2019
Earlier work this paper cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
“Learning to summarize with human feedback”
Nisan Stiennon et al · 2020
Earlier work this paper cites.
Remi Denton et al · 2021
Earlier work this paper cites.
“Distortion in social choice problems: The first 15 years and beyond”
Elliot Anshelevich, Aris Filos-Ratsikas, Nisarg Shah and Alexandros Voudouris · 2021
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“RL with KL penalties is better viewed as Bayesian inference”
Tomasz Korbak, Ethan Perez and Christopher Buckley · 2022
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward Hu et al · 2022
Earlier work this paper cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“On the identifiability of mixtures of ranking models”
Xiaomin Zhang, Xucheng Zhang, Po-Ling Loh and Yingyu Liang · 2022
Earlier work this paper cites.
“Fine-tuning language models to find agreement among humans with diverse preferences”
Michiel Bakker et al · 2022
Earlier work this paper cites.
“Why Don‘t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks”
Marta Sandri, Elisa Leonardelli, Sara Tonelli and Elisabetta Jezek · 2023
Earlier work this paper cites.
“A density estimation perspective on learning from pairwise human preferences”
Vincent Dumoulin et al · 2023
Earlier work this paper cites.
“Distributional preference learning: Understanding and accounting for hidden context in RLHF”
Anand Siththaranjan, Cassidy Laidlaw and Dylan Hadfield-Menell · 2023
Earlier work this paper cites.
“SLiC-HF: Sequence likelihood calibration with human feedback”
Yao Zhao et al · 2023
Cited alongside, same era.
“Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback” Survey Certification, Featured Certification
Stephen Casper et al · 2023
Cited alongside, same era.
“Personalized soups: Personalized large language model alignment via post-hoc parameter merging”
Joel Jang et al · 2023
Cited alongside, same era.
“Whose opinions do language models reflect?”
Shibani Santurkar et al · 2023
Cited alongside, same era.
“Diverging Preferences: When do Annotators Disagree and do Models Know?”
Michael Zhang et al · 2024
Cited alongside, same era.
“Generalized Preference Optimization: A Unified Approach to Offline Alignment”
Yunhao Tang et al · 2024
Later among the works it cites.
“SimPO: Simple Preference Optimization with a Reference-Free Reward”
Yu Meng, Mengzhou Xia and Danqi Chen · 2024
Later among the works it cites.
“Personalized language modeling from personalized human feedback”
Xinyu Li, Zachary Lipton and Liu Leqi · 2024
Later among the works it cites.
“Aligning to Thousands of Preferences via System Message Generalization”
Seongyun Lee, Sue Park, Seungone Kim and Minjoon Seo · 2024
Later among the works it cites.
“Personalized Adaptation via In-Context Preference Learning”
Allison Lau et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2024
Cited alongside, same era.
“Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback”
Vincent Conitzer et al · 2024
Cited alongside, same era.
“Axioms for AI Alignment from Human Feedback”
Luise Ge et al · 2024
Cited alongside, same era.
“Position: A Roadmap to Pluralistic Alignment”
Taylor Sorensen et al · 2024
Cited alongside, same era.
“RLHF from heterogeneous feedback via personalization and preference aggregation”
Chanwoo Park et al · 2024
Cited alongside, same era.
“Direct language model alignment from online AI feedback”
Shangmin Guo et al · 2024
Cited alongside, same era.
“The benefits, risks and bounds of personalizing the alignment of large language models to individuals”
Hannah Kirk, Bertie Vidgen, Paul Röttger and Scott Hale · 2024
Cited alongside, same era.
Tianyi Qiu · 2024
Later among the works it cites.
“Policy aggregation”
Parand Alamdari, Soroush Ebadian and Ariel Procaccia · 2024
Later among the works it cites.
“Mapping social choice theory to RLHF”
Jessica Dai and Eve Fleisig · 2024
Later among the works it cites.
“A Minimaximalist Approach to Reinforcement Learning from Human Feedback”
Gokul Swamy et al · 2024
Later among the works it cites.
Hadassah Harland et al · 2024
Later among the works it cites.
Haoxiang Wang et al · 2024
Later among the works it cites.
“Provable multi-party reinforcement learning with diverse human feedback”
Huiying Zhong et al · 2024
Later among the works it cites.
“Aligning crowd feedback via distributional preference reward modeling”
Dexun Li et al · 2024
Later among the works it cites.
“Pareto-optimal learning from preferences with hidden context”
Ryan Boldi, Li Ding, Lee Spector and Scott Niekum · 2024
Later among the works it cites.
“Rewarded soups: Towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards”
Alexandre Rame et al · 2024
Later among the works it cites.
“Preference Learning Algorithms Do Not Learn Preference Rankings”
Angelica Chen et al · 2024
Later among the works it cites.
“On Diversified Preferences of Large Language Model Alignment”, 2024
Dun Zeng et al · 2024
Later among the works it cites.
“Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models”
Hritik Bansal, John Dang and Aditya Grover · 2024
Later among the works it cites.
“Can Language Models Reason about Individualistic Human Values and Preferences?”
Liwei Jiang, Taylor Sorensen, Sydney Levine and Yejin Choi · 2024
Later among the works it cites.
“PersonalLLM: Tailoring LLMs to Individual Preferences”, 2024
Thomas. Zollo et al · 2024
Later among the works it cites.
Ariel. Procaccia, Benjamin Schiffer and Shirley Zhang · 2025
Closest in time.
“About Pew Research Center” Accessed: 2025-01-20, 2025
Pew Research Center · 2025
Closest in time.
Nishant Balepur et al · 2025
Closest in time.
“Personalized Preference Fine-tuning of Diffusion Models”
Meihua Dang et al · 2025
Closest in time.