Fetching the paper…
Reading the bibliography…
Understanding human preferences is crucial for improving foundation models and building personalized AI systems.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Rank analysis of incomplete block designs: A method of paired comparisons employing unequal repetitions on pairs
Otto Dykstra. 1960 · 1960
Earlier work this paper cites.
Response surface fitting using a generalization of the bradley-terry paired comparison model
A Springall. 1973 · 1973
Earlier work this paper cites.
Random indexing of text samples for latent semantic analysis
Pentii Kanerva, Jan Kristoferson, and Anders Holst. 2000 · 2000
Earlier work this paper cites.
Incremental learning for robust visual tracking
David A Ross, Jongwoo Lim, Ruei-Sung Lin, and Ming-Hsuan Yang. 2008 · 2008
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare. 2019 · 2019
Earlier work this paper cites.
Asymptotic theory of sparse bradley–terry model
Ruijian Han, Rougang Ye, Chunxi Tan, and Kani Chen. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Earlier work this paper cites.
Exploring dimensionality reduction techniques in multilingual transformers
Álvaro Huertas-García, Alejandro Martín, Javier Huertas-Tato, and David Camacho. 2023 · 2023
Earlier work this paper cites.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023 · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023 · 2023
Cited alongside, same era.
Epistemic neural networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy. 2023 · 2023
Cited alongside, same era.
Query-dependent prompt evaluation and optimization with offline inverse rl
Hao Sun, Alihan Hüyük, and Mihaela van der Schaar. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023 · 2023
Cited alongside, same era.
Traversing pareto optimal policies: Provably efficient multi-objective reinforcement learning
Shuang Qiu, Dake Zhang, Rui Yang, Boxiang Lyu, and Tong Zhang. 2024 · 2024
Later among the works it cites.
Dmoerm: Recipes of mixture-of-experts for effective reward modeling
Shanghaoran Quan. 2024 · 2024
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord. 2024 · 2024
Later among the works it cites.
Hao Sun, Yunyi Shen, and Jean-Francois Ton. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuanhe Tian, Ruyi Gan, Yan Song, Jiaxing Zhang, and Yongdong Zhang. 2023 · 2023
Cited alongside, same era.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, et al. 2023 · 2023
Cited alongside, same era.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Scalable ensembling for mitigating reward overoptimisation
Ahmed M Ahmed, Rafael Rafailov, Stepan Sharkov, Xuechen Li, and Sanmi Koyejo. 2024 · 2024
Cited alongside, same era.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Bedi, and Mengdi Wang. 2024 · 2024
Cited alongside, same era.
Direct preference optimization with unobserved preference heterogeneity
Keertana Chidambaram, Karthik Vinay Seetharaman, and Vasilis Syrgkanis. 2024 · 2024
Cited alongside, same era.
Rlhf workflow: From reward modeling to online rlhf
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. 2024 · 2024
Cited alongside, same era.
Uncovering latent human wellbeing in language model embeddings
Pedro Freire, ChengCheng Tan, Adam Gleave, Dan Hendrycks, and Scott Emmons. 2024 · 2024
Cited alongside, same era.
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. 2024 · 2024
Later among the works it cites.
Embedding-aligned language models
Guy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Lior Shani, Ethan Liang, and Craig Boutilier. 2024 · 2024
Later among the works it cites.
Imitating language via scalable inverse reinforcement learning
Markus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja, Jorg Bornschein, Sandy Huang, Artem Sokolov, Matt Barnes, Guillaume Desjardins, Alex Bewley, et al. 2024 · 2024
Later among the works it cites.
Teng Xiao, Mingxiao Li, Yige Yuan, Huaisheng Zhu, Chao Cui, and Vasant G Honavar. 2024 · 2024
Later among the works it cites.
Expel: Llm agents are experiential learners
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024 · 2024
Later among the works it cites.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization
Zhanhui Zhou, Jie Liu, Jing Shao, Xiangyu Yue, Chao Yang, Wanli Ouyang, and Yu Qiao. 2024 · 2024
Later among the works it cites.
Benchmarking large language models on answering and explaining challenging medical questions
Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. 2025 · 2025
Closest in time.
Pilaf: Optimal human preference sampling for reward modeling
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe, and Yaqi Duan. 2025 · 2025
Closest in time.
Reviving the classics: Active reward modeling in large language model alignment
Yunyi Shen, Hao Sun, and Jean-François Ton. 2025 · 2025
Closest in time.
Hao Sun, Yunyi Shen, Jean-Francois Ton, and Mihaela van der Schaar. 2025 · 2025
Closest in time.
Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, et al. 2025 · 2025
Closest in time.