Fetching the paper…
Reading the bibliography…
We study the tendency of AI systems to deceive by constructing a realistic simulation setting of a company AI assistant.
The elephant in the brain: Hidden motives in everyday life
Kevin Simler and Robin Hanson · 2017
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Earlier work this paper cites.
The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
Joel Lehman, Jeff Clune, Dusan Misevic, and et al · 2020
Earlier work this paper cites.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 2021
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
Joe Carlsmith · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al · 2023
Cited alongside, same era.
Hoodwinked: Deception and cooperation in a text-based game for language models
Aidan O’Gara · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Do the rewards justify the means? Measuring trade-offs between rewards and ethical behavior in the Machiavelli benchmark
AI deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Later among the works it cites.
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Later among the works it cites.
Sleeper Agents: Training deceptive LLMs that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Cited alongside, same era.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.