Fetching the paper…
Reading the bibliography…
Large language models (LLMs) trained with Reinforcement Learning from Human Feedback (RLHF) have demonstrated remarkable capabilities, but their underlying reward functions and decision-making processes remain opaque.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum margin planning
Nathan D. Ratliff, J. Andrew Bagnell, and Martin A. Zinkevich · 2006
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, et al · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, et al · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models, 2021
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, et al · 2021
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Earlier work this paper cites.
A mathematical framework for transformer circuits, 2022
Nelson Elhage, Neel Nanda, Catherine Olsson, et al · 2022
Earlier work this paper cites.
Teacher forcing recovers reward functions for text generation
Yongchang Hao, Yuxin Liu, and Lili Mou · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
Catherine Olsson, Neel Nanda, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Cited alongside, same era.
Max-margin contrastive learning
Anshul Shah, Suvrit Sra, Rama Chellappa, and Anoop Cherian · 2022
Cited alongside, same era.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song · 2023
Later among the works it cites.
Ai transparency in the age of llms: A human-centered research roadmap
Q Vera Liao and Jennifer Wortman Vaughan · 2023
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Causal confusion and reward misidentification in preference-based reward learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D Dragan, and Daniel S Brown · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Cited alongside, same era.
Stealthy and persistent unalignment on large language models via backdoor injections
Yuanpu Cao, Bochuan Cao, and Jinghui Chen · 2023
Cited alongside, same era.
Open problems in rlhf, 2023
Joshua Casper, Alice Chan, Ethan Chi, et al · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation models, 2023
Zihan Dong, Mengdi Wang, Hang Zhang, et al · 2023
Cited alongside, same era.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Later among the works it cites.
Offline prompt evaluation and optimization with inverse reinforcement learning
Hao Sun · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Later among the works it cites.
Rrhf: Rank responses to align human feedback, 2023
Weijie Yuan, Xiang Rao, Ming Wang, et al · 2023
Later among the works it cites.
Slic: Self-supervised learning of implicit contrastive objectives for language model alignment, 2023
Yuchen Zhao, Bill Yuchen Lin, et al · 2023
Later among the works it cites.
General preference optimization, 2024
Mohammad Gheshlaghi Azar, Arthur Guez, Aslan Glaese, and David Silver · 2024
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2024
Closest in time.
Supervised fine-tuning as inverse reinforcement learning, 2024
Shuai Sun, Aslan Glaese, David Krueger, et al · 2024
Closest in time.