Fetching the paper…
Reading the bibliography…
Great success has been reported using Reinforcement Learning from Human Feedback (RLHF) to align large language models, with open preference datasets enabling wider experimentation, particularly for "helpfulness" in tasks like dialogue and web question answering.
The Curious Case of Neural Text Degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 1904
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W. Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Controlling output length in neural encoder-decoders
Yuta Kikuchi, Graham Neubig, Ryohei Sasano, Hiroya Takamura, and Manabu Okumura · 2016
Earlier work this paper cites.
Controlling linguistic style aspects in neural language generation
Jessica Ficler and Yoav Goldberg · 2017
Earlier work this paper cites.
Event representations for automated story generation with deep neural nets
Lara J. Martin, Prithviraj Ammanabrolu, William Hancock, Shruti Singh, Brent Harrison, and Mark O. Riedl · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Playing text-adventure games with graph-based deep reinforcement learning
Prithviraj Ammanabrolu and Mark O. Riedl · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith · 2018
Earlier work this paper cites.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang · 2018
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
ELI5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 2019
Cited alongside, same era.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan J. Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi · 2020
Cited alongside, same era.
TRL: Transformer Reinforcement Learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, and Shengyi Huang · 2020
Cited alongside, same era.
Defining and characterizing reward hacking
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Aligning robot and human representations
Andreea Bobu, Andi Peng, Pulkit Agrawal, Julie A. Shah, and Anca D. Dragan · 2023
Closest in time.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback, 2023
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Human-centered loss functions (halos)
Kawin Ethayaraj, Winnie Xu, Dan Jurafsky, and Douwe Kiela · 2023
Closest in time.
Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, S. Arun Balaji, Jeff Wu, Ouyang Long, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2021
Cited alongside, same era.
Simon Zhuang and Dylan Hadfield-Menell · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, T. J. Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan · 2022
Cited alongside, same era.
On length divergence bias in textual matching models
Lan Jiang, Tianshu Lyu, Yankai Lin, Meng Chong, Xiaoyong Lyu, and Dawei Yin · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe · 2022
Cited alongside, same era.
Reward gaming in conditional text generation
Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, and He He · 2022
Cited alongside, same era.
Closest in time.
HuggingFace H4 Stack Exchange Preference Dataset, 2023
Nathan Lambert, Lewis Tunstall, Nazneen Rajani, and Tristan Thrush · 2023
Closest in time.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Closest in time.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2023
Closest in time.
Simeng Sun, Dhawal Gupta, and Mohit Iyyer · 2023
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hanna Hajishirzi · 2023
Closest in time.
RLCD: Reinforcement Learning from Contrast Distillation for Language Model Alignment
Kevin Yang, Dan Klein, Asli Celikyilmaz, Nanyun Peng, and Yuandong Tian · 2023
Closest in time.
SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu · 2023
Closest in time.