Fetching the paper…
Reading the bibliography…
In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals.
Risks from learned optimization in advanced machine learning systems, 2019
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Probable inference, the law of succession, and statistical inference
Edwin B Wilson · 1927
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Understanding the failure modes of out-of-distribution generalization
Vaishnavh Nagarajan, Anders Andreassen, and Behnam Neyshabur · 2010
Earlier work this paper cites.
Comparison of confidence intervals for the Poisson mean: Some new aspects
V. V. Patil and H. V. Kulkarni · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, D. Erhan, Ian J. Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Faulty reward functions in the wild, 12 2016
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier, 2016
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search, 2017
Thomas Anthony, Zheng Tian, and David Barber · 2017
Earlier work this paper cites.
Adversarial examples in the physical world, 2017
Alexey Kurakin, Ian Goodfellow, and Samy Bengio · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Scaling provable adversarial defenses, 2018
Eric Wong, Frank R. Schmidt, Jan Hendrik Metzen, and J. Zico Kolter · 2018
Earlier work this paper cites.
How useful is quantilization for mitigating specification-gaming?
Ryan Carey · 2019
Cited alongside, same era.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann · 2020
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity, April 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
Manipulating the distributions of experience used for self-play learning in expert iteration, 2020
Dennis J. N. J. Soemers, Éric Piette, Matthew Stephenson, and Cameron Browne · 2020
Cited alongside, same era.
Avoiding tampering incentives in deep rl via decoupled approval
Jonathan Uesato, Ramana Kumar, Victoria Krakovna, Tom Everitt, Richard Ngo, and Shane Legg · 2020
Cited alongside, same era.
Active learning helps pretrained models learn the intended task, 2022
Alex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu, and Noah Goodman · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models, 2022
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Understanding strategic deception and deceptive alignment, 9 2023
Apollo Research · 2023
Later among the works it cites.
Taken out of context: On measuring situational awareness in llms, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Later among the works it cites.
Poisoning web-scale training datasets is practical
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A general language assistant as a laboratory for alignment, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan · 2021
Cited alongside, same era.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover, 2021
Ajeya Cotra · 2021
Cited alongside, same era.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective, 2021
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 2021
Cited alongside, same era.
User tampering in reinforcement learning recommender systems
Atoosa Kasirzadeh and Charles Evans · 2021
Cited alongside, same era.
Objective robustness in deep reinforcement learning, 05 2021
Jack Koch, Lauro Langosco, Jacob Pfau, James Le, and Lee Sharkey · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Cited alongside, same era.
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback, 2023
Javier Rando and Florian Tramèr · 2023
Later among the works it cites.
Towards understanding sycophancy in language models, 2023
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2023
Later among the works it cites.
On the exploitability of instruction tuning, 2023
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Ai in software engineering at google: Progress and the path ahead, June 2024
Satish Chandra and Maxim Tabachnyk · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Instrumental deception and manipulation in llms - a case study
Olli Järviniemi · 2024
Closest in time.
Reward hacking behavior can generalize across tasks
Kei Nishimura-Gasparian, Isaac Dunn, Henry Sleight, Miles Turpin, Evan Hubinger, Carson Denison, and Ethan Perez · 2024
Closest in time.
Inducing unprompted misalignment in llms
Sam Svenningsen, Evan Hubinger, and Henry Sleight · 2024
Closest in time.