Fetching the paper…
Reading the bibliography…
We introduce a method to address goal misgeneralization in reinforcement learning (RL), leveraging Large Language Model (LLM) feedback during training.
Superintelligence: Paths, dangers, strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Concrete problems in ai safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Learning from human preferences, openai, 2017
Dario Amodei, Paul Christiano, and Alex Ray · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks, 2017
Andre Esteva, Brett Kuprel, , Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel, 2017
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Ai safety gridworlds, 2017
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures, 2018
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari, 2018
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Ai safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Constructing unrestricted adversarial examples with generative models, 2018
Yang Song, Rui Shu, Nate Kushman, and Stefano Ermon · 2018
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning, 2019
Oriol Vinyals, Igor Babuschkin, Michaël Mathieu Wojciech M Czarnecki, Junyoung Chung Andrew Dudzik, David H Choi, Richard Powell, Timo Ewalds, and et al. Petko Georgiev · 2019
Cited alongside, same era.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Cited alongside, same era.
An overview of 11 proposals for building safe advanced ai, 2020
Evan Hubinger · 2020
Cited alongside, same era.
Training procgen environment with py- torch
Hojoon Lee · 2020
Cited alongside, same era.
Fine-tuning language models from human preferences, 2020
When life gives you lemons, make cherryade: Converting feedback from bad responses into good labels, 2022
Weiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, and Jing Xu · 2022
Later among the works it cites.
Defining and characterizing reward gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Learning to summarize from human feedback, 2022
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability, 2022
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, Buck Shlegeris, and Nate Thomas · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold, 2021
John Jumper, Richard Evans, Alexander Pritzel, Michael Figurnov Tim Green, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis · 2021
Cited alongside, same era.
Ethical-advice taker: Do language models understand natural language interventions?, 2021
Jieyu Zhao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Kai-Wei Chang · 2021
Cited alongside, same era.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models, 2022
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Cited alongside, same era.
Red teaming language models with language models, 2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators, 2022
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Cited alongside, same era.
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah · 2023
Later among the works it cites.
Harms from increasingly agentic algorithmic systems
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstantinos Voudouris, Umang Bhatt, Adrian Weller, David Krueger, and Tegan Maharaj · 2023
Later among the works it cites.
Aligning ai with shared human values, 2023
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback, 2023
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2023
Later among the works it cites.
Goal misgeneralization in deep reinforcement learning, 2023
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, Laurent Orseau, and David Krueger · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback, 2023
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Later among the works it cites.
Large language models sensitivity to the order of options in multiple-choice questions., 2023
Pouya Pezeshkpour and Estevam Hruschka · 2023
Later among the works it cites.