Fetching the paper…
Reading the bibliography…
A key challenge in training Large Language Models (LLMs) is properly aligning them with human preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Independence of clones as a criterion for voting rules
T. N. Tideman · 1987
Earlier work this paper cites.
Differential privacy
C. Dwork · 2006
Earlier work this paper cites.
Cloning in elections
E. Elkind, P. Faliszewski, and A. Slinko · 2010
Earlier work this paper cites.
Clone structures in voters’ preferences
E. Elkind, P. Faliszewski, and A. Slinko · 2012
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Obvious independence of clones
R. E. Berker, S. Casacuberta, C. Ong, and I. Robinson · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
On the identifiability of mixtures of ranking models
X. Zhang, X. Zhang, P.-L. Loh, and Y. Liang · 2022
Earlier work this paper cites.
Large language models as simulated economic agents: What can we learn from homo silicus?
J. J. Horton · 2023
Cited alongside, same era.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
A. Siththaranjan, C. Laidlaw, and D. Hadfield-Menell · 2023
Cited alongside, same era.
Rlhf and iia: Perverse incentives
W. Xu, S. Dong, X. Lu, G. Lam, Z. Wen, and B. Van Roy · 2023
Cited alongside, same era.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
B. Zhu, M. Jordan, and J. Jiao · 2023
Cited alongside, same era.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2023
Mapping social choice theory to rlhf
J. Dai and E. Fleisig · 2024
Later among the works it cites.
Axioms for ai alignment from human feedback
L. Ge, D. Halpern, E. Micha, A. D. Procaccia, I. Shapira, Y. Vorobeychik, and J. Wu · 2024
Later among the works it cites.
Corruption robust offline reinforcement learning with human feedback
D. Mandal, A. Nika, P. Kamalaruban, A. Singla, and G. Radanović · 2024
Later among the works it cites.
Rlhf from heterogeneous feedback via personalization and preference aggregation
C. Park, M. Liu, D. Kong, K. Zhang, and A. E. Ozdaglar · 2024
Later among the works it cites.
Personalizing reinforcement learning from human feedback with variational preference learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Foundational challenges in assuring alignment and safety of large language models
U. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, E. S. Lubana, E. Jenner, S. Casper, O. Sourbut, et al · 2024
Cited alongside, same era.
Robust reinforcement learning from corrupted human feedback
A. Bukharin, I. Hong, H. Jiang, Z. Li, Q. Zhang, Z. Zhang, and T. Zhao · 2024
Cited alongside, same era.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, F. Huang, D. Manocha, A. S. Bedi, and M. Wang · 2024
Cited alongside, same era.
Social choice should guide ai alignment in dealing with diverse human feedback
V. Conitzer, R. Freedman, J. Heitzig, W. H. Holliday, B. M. Jacobs, N. Lambert, M. Mossé, E. Pacuit, S. Russell, H. Schoelkopf, et al · 2024
Cited alongside, same era.
S. Poddar, Y. Wan, H. Ivison, A. Gupta, and N. Jaques · 2024
Later among the works it cites.
A roadmap to pluralistic alignment
T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, et al · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
G. Swamy, C. Dann, R. Kidambi, Z. S. Wu, and A. Agarwal · 2024
Later among the works it cites.
Provable multi-party reinforcement learning with diverse human feedback
H. Zhong, Z. Deng, W. J. Su, Z. S. Wu, and L. Zhang · 2024
Later among the works it cites.