Fetching the paper…
Reading the bibliography…
Large foundation models pretrained on raw web-scale data are not readily deployable without additional step of extensive alignment to human preferences.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 1908
Earlier work this paper cites.
Psychological scaling without a unit of measurement
Clyde H Coombs · 1950
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 1959
Earlier work this paper cites.
Metric structures in ordinal data
Roger N Shepard · 1966
Earlier work this paper cites.
Generalized linear models
J. A. Nelder and R. W. M. Wedderburn · 1972
Earlier work this paper cites.
Ideal point models of preference
Joel Huber · 1976
Earlier work this paper cites.
Choosing preferences by constructing institutions: A cultural theory of preference formation
Aaron Wildavsky · 1987
Earlier work this paper cites.
Learning a distance metric from relative comparisons
Matthew Schultz and Thorsten Joachims · 2003
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
David R Hunter · 2004
Earlier work this paper cites.
How to rank with few errors
Claire Kenyon-Mathieu and Warren Schudy · 2007
Earlier work this paper cites.
Noisy sorting without resampling
Mark Braverman and Elchanan Mossel · 2007
Earlier work this paper cites.
Human preference for individual colors
Stephen E Palmer and Karen B Schloss · 2010
Earlier work this paper cites.
Preference learning and ranking by pairwise comparison
Johannes Fürnkranz and Eyke Hüllermeier · 2010
Earlier work this paper cites.
Active ranking using pairwise comparisons
Kevin G Jamieson and Robert Nowak · 2011
Earlier work this paper cites.
Adaptively learning the crowd kernel
Omer Tamuz, Ce Liu, Serge Belongie, Ohad Shamir, and Adam Tauman Kalai · 2011
Earlier work this paper cites.
Iterative ranking from pair-wise comparisons
Sahand Negahban, Sewoong Oh, and Devavrat Shah · 2012
Earlier work this paper cites.
Metric learning: A survey
Brian Kulis · 2013
Earlier work this paper cites.
Learning to top-k search using pairwise comparisons
Brian Eriksson · 2013
Earlier work this paper cites.
A statistical convergence perspective of algorithms for rank aggregation from pairwise data
Arun Rajkumar and Shivani Agarwal · 2014
Earlier work this paper cites.
Uniqueness of ordinal embedding
Matthäus Kleindessner and Ulrike Luxburg · 2014
Earlier work this paper cites.
Metric learning
Aurélien Bellet, Amaury Habrard, and Marc Sebban · 2015
Earlier work this paper cites.
Robustness and generalization for metric learning
Aurélien Bellet and Amaury Habrard · 2015
Earlier work this paper cites.
Evaluating change in behavioral preferences: Multidimensional scaling single-ideal point model
Cody Ding · 2016
Earlier work this paper cites.
Actively learning hemimetrics with applications to eliciting user preferences
Adish Singla, Sebastian Tschiatschek, and Andreas Krause · 2016
Earlier work this paper cites.
Stochastically transitive models for pairwise comparisons: Statistical and computational issues
Nihar Shah, Sivaraman Balakrishnan, Aditya Guntuboyina, and Martin Wainwright · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Simple, robust and optimal ranking from pairwise comparisons
Nihar B Shah and Martin J Wainwright · 2017
Cited alongside, same era.
Learning low-dimensional metrics
Blake Mason, Lalit Jain, and Robert Nowak · 2017
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
Neuroaesthetics and art’s diversity and universality
Marcos Nadal and Anjan Chatterjee · 2019
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Scaling up gans for text-to-image synthesis
Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park, Eli Shechtman, Sylvain Paris, and Taesung Park · 2023
Later among the works it cites.
Stylegan-t: Unlocking the power of gans for fast large-scale text-to-image synthesis
Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger, and Timo Aila · 2023
Later among the works it cites.
Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jung, Khyathi Chandu, Abhilasha Ravichander, Prithviraj Ammanabrolu, Liwei Jiang, Sahana Ramnath, Nouha Dziri, et al · 2023
Later among the works it cites.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simultaneous preference and metric learning from paired comparisons
Austin Xu and Mark Davenport · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Cited alongside, same era.
Later among the works it cites.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models, 2023
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Later among the works it cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Later among the works it cites.
Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al · 2023
Later among the works it cites.
Rewarding chatbots for real-world engagement with millions of users
Robert Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Ziyi Zhu, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Liu, Fritz Cremer, Valentin Assassi, et al · 2023
Later among the works it cites.
Ai alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al · 2023
Later among the works it cites.
Towards measuring the representation of subjective global opinions in language models, 2024
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli · 2024
Closest in time.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Closest in time.
Hyeong Kyu Choi and Yixuan Li · 2024
Closest in time.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Rame, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, and Matthieu Cord · 2024
Closest in time.
Personalized language modeling from personalized human feedback
Xinyu Li, Zachary C Lipton, and Liu Leqi · 2024
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2024
Closest in time.