Fetching the paper…
Reading the bibliography…
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models with human values, reward hacking, also termed reward overoptimization, remains a critical challenge.
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, DJ Strouse, Zafarali Ahmed, Hugo Larochelle, Matthew Botvinick, Yoshua Bengio, and Sergey Levine · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 1912
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, Xiaowei Xu, et al · 1996
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
An experimental comparison of click position-bias models
Nick Craswell, Onno Zoeter, Michael Taylor, and Bill Ramsey · 2008
Earlier work this paper cites.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2012
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Earlier work this paper cites.
Compressing neural networks using the variational information bottleneck
Bin Dai, Chen Zhu, Baining Guo, and David Wipf · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Position bias estimation for unbiased learning to rank in personal search
Xuanhui Wang, Nadav Golbandi, Michael Bendersky, Donald Metzler, and Marc Najork · 2018
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker · 2019
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Mkqa: A linguistically diverse benchmark for multilingual open domain question answering
Shayne Longpre, Yi Lu, and Joachim Daiber · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al · 2023
Later among the works it cites.
Confronting reward model overoptimization with constrained rlhf
Ted Moskovitz, Aaditya K Singh, DJ Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca D Dragan, and Stephen McAleer · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah · 2023
Later among the works it cites.
Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback
Wei Shen, Rui Zheng, Wenyu Zhan, Jun Zhao, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
Information-theoretic odometry learning
Sen Zhang, Jing Zhang, and Dacheng Tao · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Cited alongside, same era.
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, et al · 2023
Later among the works it cites.
Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles
Yuanzhao Zhai, Han Zhang, Yu Lei, Yue Yu, Kele Xu, Dawei Feng, Bo Ding, and Huaimin Wang · 2023
Later among the works it cites.
Delve into PPO: Implementation matters for stable RLHF
Rui Zheng, Shihan Dou, Songyang Gao, Yuan Hua, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Yuhao Zhou, Limao Xiong, Lu Chen, Zhiheng Xi, Nuo Xu, Wenbin Lai, Minghao Zhu, Haoran Huang, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Later among the works it cites.
ODIN: Disentangled reward mitigates hacking in rlhf
Lichang Chen, Chen Zhu, Davit Soselia, Jiuhai Chen, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Closest in time.
Reward model ensembles help mitigate overoptimization
Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al · 2024
Closest in time.
Warm: On the benefits of weight averaged reward models
Alexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi, Geoffrey Cideron, Olivier Bachem, and Johan Ferret · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
Binghai Wang, Rui Zheng, Lu Chen, Yan Liu, Shihan Dou, Caishuang Huang, Wei Shen, Senjie Jin, Enyu Zhou, Chenyu Shi, et al · 2024
Closest in time.
Improving generalization of alignment with human preferences through group invariant learning
Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, et al · 2024
Closest in time.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf
Banghua Zhu, Michael I Jordan, and Jiantao Jiao · 2024
Closest in time.