Fetching the paper…
Reading the bibliography…
In practice, preference learning from human feedback depends on incomplete data with hidden context.
Fine-Tuning Language Models from Human Preferences, January 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing, July 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 1910
Earlier work this paper cites.
Incentive Compatible Active Learning, November 2019
Federico Echenique and Siddharth Prasad · 1911
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
A Statistical Convergence Perspective of Algorithms for Rank Aggregation from Pairwise Data
Arun Rajkumar and Shivani Agarwal · 1938
Earlier work this paper cites.
The Borda count and agenda manipulation
Michael Dummett · 1998
Earlier work this paper cites.
Comparative analysis of Bradley-Terry and Thurstone-Mosteller paired comparison models for image quality assessment
John C. Handley · 2001
Earlier work this paper cites.
Reward-Rational (Implicit) Choice: A Unifying Formalism for Reward Learning
Hong Jun Jeon, Smitha Milli, and Anca D. Dragan · 2002
Earlier work this paper cites.
Common Voting Rules as Maximum Likelihood Estimators
Vincent Conitzer and Tuomas Sandholm · 2005
Earlier work this paper cites.
Voting systems
Paul E. Johnson · 2005
Earlier work this paper cites.
Multi-Principal Assistance Games, July 2020
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell, and Stuart Russell · 2007
Earlier work this paper cites.
Efficient Bayesian Inference for Generalized Bradley-Terry Models, November 2010
Francois Caron and Arnaud Doucet · 2010
Earlier work this paper cites.
APRIL: Active Preference-learning based Reinforcement Learning, August 2012
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
Math in Society
David Lippman · 2012
Earlier work this paper cites.
The original Borda count and partial voting
Peter Emerson · 2013
Earlier work this paper cites.
A Survey of Preference-Based Online Learning with Bandit Algorithms
Róbert Busa-Fekete and Eyke Hüllermeier · 2014
Earlier work this paper cites.
Spectral MLE: Top-$K$ Rank Aggregation from Pairwise Comparisons, May 2015
Yuxin Chen and Changho Suh · 2015
Earlier work this paper cites.
Fast and Accurate Inference of Plackett– Luce Models
Lucas Maystre and Matthias Grossglauser · 2015
Cited alongside, same era.
Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence, May 2015
Nihar B. Shah, Sivaraman Balakrishnan, Joseph Bradley, Abhay Parekh, Kannan Ramchandran, and Martin J. Wainwright · 2015
Cited alongside, same era.
Deep reinforcement learning from human preferences, June 2017
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Active Preference-Based Learning of Reward Functions
Dorsa Sadigh, Anca Dragan, Shankar Sastry, and Sanjit Seshia · 2017
Cited alongside, same era.
Approximate Ranking from Pairwise Comparisons, January 2018
Reinhard Heckel, Max Simchowitz, Kannan Ramchandran, and Martin J. Wainwright · 2018
Dueling RL: Reinforcement Learning with Trajectory Preferences, November 2021
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2021
Later among the works it cites.
Models of human preference for learning reward functions, June 2022
W. Bradley Knox, Stephane Hatgis-Kessell, Serena Booth, Scott Niekum, Peter Stone, and Alessandro Allievi · 2022
Later among the works it cites.
Cassidy Laidlaw and Anca Dragan · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simple, Robust and Optimal Ranking from Pairwise Comparisons
Nihar B. Shah and Martin J. Wainwright · 2018
Cited alongside, same era.
Decoupled Weight Decay Regularization, January 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
LESS is More: Rethinking Probabilistic Models of Human Behavior
Andreea Bobu, Dexter R. R. Scobee, Jaime F. Fisac, S. Shankar Sastry, and Anca D. Dragan · 2020
Cited alongside, same era.
Minimax Rate for Learning From Pairwise Comparisons in the BTL Model
Julien Hendrickx, Alex Olshevsky, and Venkatesh Saligrama · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Learning Mixtures of Plackett-Luce Models, March 2020
Zhibing Zhao, Peter Piech, and Lirong Xia · 2020
Cited alongside, same era.
Consequences of Misaligned AI
Simon Zhuang and Dylan Hadfield-Menell · 2020
Cited alongside, same era.
Which Examples Should be Multiply Annotated? Active Learning When Annotators May Disagree
Connor Baumler, Anna Sotnikova, and Hal Daumé III · 2023
Closest in time.
Safe RLHF: Safe Reinforcement Learning from Human Feedback, October 2023
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback, August 2023
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
A density estimation perspective on learning from pairwise human preferences, 2023
Vincent Dumoulin, Daniel D. Johnson, Pablo Samuel Castro, Hugo Larochelle, and Yann Dauphin · 2023
Closest in time.
Generative Social Choice, September 2023
Sara Fish, Paul Gölz, David C. Parkes, Ariel D. Procaccia, Gili Rusak, Itai Shapira, and Manuel Wüthrich · 2023
Closest in time.
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, and Dan Klein · 2023
Closest in time.
Embedding Democratic Values into Social Media AIs via Societal Objective Functions, July 2023
Chenyan Jia, Michelle S. Lam, Minh Chau Mai, Jeff Hancock, and Michael S. Bernstein · 2023
Closest in time.
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications, October 2023
Abhilash Mishra · 2023
Closest in time.
Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models, July 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Closest in time.
Jailbroken: How Does LLM Safety Training Fail?, July 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Banghua Zhu, Jiantao Jiao, and Michael I. Jordan · 2023
Closest in time.