Fetching the paper…
Reading the bibliography…
We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own preferences.
Manipulation of voting schemes: a general result
Allan Gibbard · 1973
Earlier work this paper cites.
Strategy-proofness and arrow’s conditions: Existence and correspondence theorems for voting procedures and social welfare functions
Mark Allen Satterthwaite · 1975
Earlier work this paper cites.
Straightforwardness of game forms with lotteries as outcomes
Allan Gibbard · 1978
Earlier work this paper cites.
On strategy-proofness and single peakedness
Hervé Moulin · 1980
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Earlier work this paper cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al · 2022
Earlier work this paper cites.
When are offline two-player zero-sum markov games solvable?
Qiwen Cui and Simon S Du · 2022
Earlier work this paper cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Earlier work this paper cites.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte · 2023
Earlier work this paper cites.
A survey of reinforcement learning from human feedback
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier · 2023
Earlier work this paper cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Cited alongside, same era.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Cited alongside, same era.
Robust reinforcement learning from corrupted human feedback
Alexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li, Qingru Zhang, Zixuan Zhang, and Tuo Zhao · 2024
Cited alongside, same era.
Persona: A reproducible testbed for pluralistic alignment
Louis Castricato, Nathan Lile, Rafael Rafailov, Jan-Philipp Fränken, and Chelsea Finn · 2024
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al · 2024
Later among the works it cites.
Strategic linear contextual bandits
Thomas Kleine Buening, Aadirupa Saha, Christos Dimitrakakis, and Haifeng Xu · 2024
Later among the works it cites.
Corruption robust offline reinforcement learning with human feedback
Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban, Adish Singla, and Goran Radanović · 2024
Later among the works it cites.
Rlhf from heterogeneous feedback via personalization and preference aggregation
Chanwoo Park, Mingyang Liu, Dingwen Kong, Kaiqing Zhang, and Asuman E Ozdaglar · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Maxmin-rlhf: Alignment with diverse human preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Dinesh Manocha, Furong Huang, Amrit Bedi, and Mengdi Wang · 2024
Cited alongside, same era.
Pal: Pluralistic alignment framework for learning from heterogeneous preferences
Daiwei Chen, Yi Chen, Aniket Rege, and Ramya Korlakai Vinayak · 2024
Cited alongside, same era.
Rime: Robust preference-based reinforcement learning with noisy preferences
Jie Cheng, Gang Xiong, Xingyuan Dai, Qinghai Miao, Yisheng Lv, and Fei-Yue Wang · 2024
Cited alongside, same era.
Social choice should guide ai alignment in dealing with diverse human feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al · 2024
Cited alongside, same era.
Modular pluralism: Pluralistic alignment via multi-llm collaboration
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov · 2024
Cited alongside, same era.
Axioms for ai alignment from human feedback
Luise Ge, Daniel Halpern, Evi Micha, Ariel D Procaccia, Itai Shapira, Yevgeniy Vorobeychik, and Junlin Wu · 2024
Cited alongside, same era.
Online learning from strategic human feedback in llm fine-tuning
Shugang Hao and Lingjie Duan · 2024
Cited alongside, same era.
Shyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou Ammar, and Ilija Bogunovic · 2024
Later among the works it cites.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Later among the works it cites.
Truthful aggregation of llms with an application to online advertising
Ermis Soumalias, Michael J Curry, and Sven Seuken · 2024
Later among the works it cites.
Mechanism design for llm fine-tuning with multiple reward models
Haoran Sun, Yurong Chen, Siwei Wang, Wei Chen, and Xiaotie Deng · 2024
Later among the works it cites.
Provable multi-party reinforcement learning with diverse human feedback
Huiying Zhong, Zhun Deng, Weijie J Su, Zhiwei Steven Wu, and Linjun Zhang · 2024
Later among the works it cites.
Policy aggregation
Parand A Alamdari, Soroush Ebadian, and Ariel D Procaccia · 2025
Closest in time.
Medical large language models are vulnerable to data-poisoning attacks
Daniel Alexander Alber, Zihao Yang, Anton Alyakin, Eunice Yang, Sumedha Rai, Aly A Valliani, Jeff Zhang, Gabriel R Rosenbaum, Ashley K Amend-Thomas, David B Kurland, et al · 2025
Closest in time.
Exploiting llm quantization
Kazuki Egashira, Mark Vero, Robin Staab, Jingxuan He, and Martin Vechev · 2025
Closest in time.