Fetching the paper…
Reading the bibliography…
Aligning large language models (LLMs) with human preferences through reinforcement learning (RLHF) can lead to reward hacking, where LLMs exploit failures in the reward model (RM) to achieve seemingly high rewards without meeting the underlying objectives.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Bounded rationality
Herbert A Simon · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Ron Kohavi, David H Wolpert, et al · 1996
Earlier work this paper cites.
Generalization error of ensemble estimators
Naonori Ueda and Ryohei Nakano · 1996
Earlier work this paper cites.
Improving ratings: audit in the british university system
Marilyn Strathern · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al · 2000
Earlier work this paper cites.
Reinforcement learning in feedback control: Challenges and benchmarks from technical process control
Roland Hafner and Martin Riedmiller · 2011
Earlier work this paper cites.
Domain generalization via invariant feature representation
Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2014
Earlier work this paper cites.
Policy gradient in lipschitz markov decision processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Faulty Reward Functions in the Wild
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
Alignment for advanced machine learning systems
Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, and Andrew Critch · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Deal or no deal? end-to-end learning for negotiation dialogues
Mike Lewis, Denis Yarats, Yann N Dauphin, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E Turner, and Douglas Eck · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein · 2017
Earlier work this paper cites.
Formal guarantees on the robustness of a classifier against adversarial manipulation
Matthias Hein and Maksym Andriushchenko · 2017
Earlier work this paper cites.
Robust large margin deep neural networks
Jure Sokolić, Raja Giryes, Guillermo Sapiro, and Miguel RD Rodrigues · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Robust loss functions under label noise for deep neural networks
Aritra Ghosh, Himanshu Kumar, and P Shanti Sastry · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmüller, and Owain Evans · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study
John R. Zech, Marcus A. Badgeley, Manway Liu, Anthony B. Costa, Joseph J. Titano, and Eric Karl Oermann · 2018
Earlier work this paper cites.
Improving stability in deep reinforcement learning with weight averaging
Evgenii Nikishin, Pavel Izmailov, Ben Athiwaratkun, Dmitrii Podoprikhin, Timur Garipov, Pavel Shvechikov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei · 2018
Earlier work this paper cites.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Invariant risk minimization
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Earlier work this paper cites.
On the feasibility of learning, rather than assuming, human biases for reward inference
Rohin Shah, Noah Gundotra, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Learning from noisy labels by regularized estimation of annotator confusion
Ryutaro Tanno, Ardavan Saeedi, Swami Sankaranarayanan, Daniel C Alexander, and Nathan Silberman · 2019
Earlier work this paper cites.
Ensemble learning in the presence of noise
Maryam Sabzevari · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Consequences of misaligned AI
Simon Zhuang and Dylan Hadfield-Menell · 2020
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 2020
Earlier work this paper cites.
What is being transferred in transfer learning?
Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang · 2020
Earlier work this paper cites.
Gradient starvation: A learning proclivity in neural networks
Mohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie · 2020
Earlier work this paper cites.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, et al · 2020
Cited alongside, same era.
Multi-agent communication meets natural language: Synergies between functional and structural language learning
Angeliki Lazaridou, Anna Potapenko, and Olivier Tieleman · 2020
Cited alongside, same era.
Countering language drift with seeded iterated learning
Yuchen Lu, Soumye Singhal, Florian Strub, Aaron Courville, and Olivier Pietquin · 2020
Cited alongside, same era.
Learning human objectives by evaluating hypothetical behavior
Siddharth Reddy, Anca Dragan, Sergey Levine, Shane Legg, and Jan Leike · 2020
Cited alongside, same era.
Adversarial robustness through local lipschitzness
Yao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov, and Kamalika Chaudhuri · 2020
Cited alongside, same era.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Later among the works it cites.
The alignment ceiling: Objective mismatch in reinforcement learning from human feedback
Nathan Lambert and Roberto Calandra · 2023
Later among the works it cites.
A long way to go: Investigating length correlations in rlhf
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A case for new neural network smoothness constraints
Mihaela Rosca, Theophane Weber, Arthur Gretton, and Shakir Mohamed · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Cited alongside, same era.
SWAD: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
In search of lost domain generalization
Ishaan Gulrajani and David Lopez-Paz · 2021
Cited alongside, same era.
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Later among the works it cites.
The political ideology of conversational ai: Converging evidence on chatgpt’s pro-environmental, left-libertarian orientation
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte · 2023
Later among the works it cites.
Natural selection favors AIs over humans
Dan Hendrycks · 2023
Later among the works it cites.
Benchmarks and algorithms for offline preference-based reward learning
Daniel Shin, Anca Dragan, and Daniel S. Brown · 2023
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Alexandre Ramé, Guillaume Couairon, Mustafa Shukor, Corentin Dancette, Jean-Baptiste Gaya, Laure Soulier, and Matthieu Cord · 2023
Later among the works it cites.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D’Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, et al · 2023
Later among the works it cites.
Reward model ensembles help mitigate overoptimization
Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2023
Later among the works it cites.
Model Ratatouille: Recycling diverse models for out-of-distribution generalization
Alexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord, Léon Bottou, and David Lopez-Paz · 2023
Later among the works it cites.
Fuse to forget: Bias reduction and selective memorization through model fusion
Kerem Zaman, Leshem Choshen, and Shashank Srivastava · 2023
Later among the works it cites.
RLAIF: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Later among the works it cites.
ID and OOD performance are sometimes inversely correlated on real-world datasets
Damien Teney, Yong Lin, Seong Joon Oh, and Ehsan Abbasnejad · 2023
Later among the works it cites.
On the challenges and practices of reinforcement learning from real human feedback
Timo Kaufmann, Sarah Ball, Jacob Beck, Eyke Hüllermeier, and Frauke Kreuter · 2023
Later among the works it cites.
Rewarding chatbots for real-world engagement with millions of users
Robert Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Liu, Fritz Cremer, Valentin Assassi, Christie-Carol Beauchamp, et al · 2023
Later among the works it cites.
Failure modes of learning reward models for llms and other sequence models
Silviu Pitis · 2023
Later among the works it cites.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Later among the works it cites.
State of what art? a call for multi-prompt llm evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2023
Later among the works it cites.
Secrets of rlhf in large language models part ii: Reward modeling
Binghai Wang et al · 2023
Later among the works it cites.
Specific versus general principles for constitutional ai
Sandipan Kundu, Yuntao Bai, Saurav Kadavath, Amanda Askell, Andrew Callahan, Anna Chen, Anna Goldie, Avital Balwit, Azalia Mirhoseini, Brayden McLean, et al · 2023
Later among the works it cites.
Knowledge is a region in weight space for fine-tuned language models
Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2023
Later among the works it cites.
PaLM 2 technical report
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
KL divergence of max-of-n, 2023
Jacob Hilton · 2023
Later among the works it cites.
Building Machine Learning Models Like Open Source Software
Colin Raffel · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson · 2023
Later among the works it cites.
ColD fusion: Collaborative descent for distributed multitask finetuning
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2023
Later among the works it cites.
Unival: Unified model for image, video, audio and language
Mustafa Shukor, Corentin Dancette, Alexandre Ramé, and Matthieu Cord · 2023
Later among the works it cites.
Seasoning model soups for robustness to adversarial and natural distribution shifts
Francesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, and Sven Gowal · 2023
Later among the works it cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2023
Later among the works it cites.
Merging decision transformers: Weight averaging for forming multi-task policies
Daniel Lawson and Ahmed H Qureshi · 2023
Later among the works it cites.
Language model alignment with elastic reset
Michael Noukhovitch, Samuel Lavoie, Florian Strub, and Aaron Courville · 2023
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Later among the works it cites.
Elastic weight removal for faithful and abstractive dialogue generation
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo M Ponti · 2023
Later among the works it cites.
Neftune: Noisy embeddings improve instruction finetuning
Neel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, et al · 2023
Later among the works it cites.
Learning optimal advantage from preferences and mistaking it for reward
W Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson, Serena Booth, Anca Dragan, Peter Stone, and Scott Niekum · 2023
Later among the works it cites.
Active reward learning from multiple teachers
Peter Barnett, Rachel Freedman, Justin Svegliato, and Stuart Russell · 2023
Later among the works it cites.
The impact of preference agreement in reinforcement learning from human feedback: A case study in summarization
Sian Gooding and Hassan Mansoor · 2023
Later among the works it cites.
Tool-augmented reward modeling
Lei Li, Yekun Chai, Shuohuan Wang, Yu Sun, Hao Tian, Ningyu Zhang, and Hua Wu · 2023
Later among the works it cites.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al · 2023
Later among the works it cites.
RIME: Robust preference-based reinforcement learning with noisy human preferences
Anonymous · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos · 2023
Later among the works it cites.
Vanishing gradients in reinforcement finetuning of language models
Noam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak, Arwen Bradley, Preetum Nakkiran, Joshua Susskind, and Etai Littwin · 2023
Later among the works it cites.
Spurious feature diversification improves out-of-distribution generalization
Yong Lin, Lu Tan, Yifan Hao, Honam Wong, Hanze Dong, Weizhong Zhang, Yujiu Yang, and Tong Zhang · 2024
Closest in time.
Theoretical guarantees on the best-of-n alignment policy
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh · 2024
Closest in time.
NeuralBeagle14-7B
Maxime Labonne · 2024
Closest in time.