Fetching the paper…
Reading the bibliography…
We demonstrate a situation in which Large Language Models, trained to be helpful, harmless, and honest, can display misaligned behavior and strategically deceive their users about this behavior without being instructed to do so.
Towards defining deception in structural causal games
Francis Rhys Ward · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
The taskrabbit example
ARCevals · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Earlier work this paper cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Earlier work this paper cites.
Deception abilities emerged in large language models
Thilo Hagendorff · 2023
Earlier work this paper cites.
Understanding strategic deception and deceptive alignment
Marius Hobbhahn · 2023
Earlier work this paper cites.
Evaluating language-model agents on realistic autonomous tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, and Paul Christiano · 2023
Earlier work this paper cites.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al · 2023
Cited alongside, same era.
Hoodwinked: Deception and cooperation in a text-based game for language models
Aidan O’Gara · 2023
Cited alongside, same era.
Gpt-4 technical report
R OpenAI · 2023
Cited alongside, same era.
How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions
Lorenzo Pacchiardi, Alex J Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y Pan, Yarin Gal, Owain Evans, and Jan Brauner · 2023
Cited alongside, same era.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Superagi
SuperAGI · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Closest in time.
Evaluating shutdown avoidance of language models in textual scenarios, 2023
Teun van der Weij, Simon Lermen, and Leon lang · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ai deception: A survey of examples, risks, and potential solutions
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Cited alongside, same era.
Scalable and transferable black-box jailbreaks for language models via persona modulation, 2023
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando · 2023
Cited alongside, same era.
Role-play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al · 2023
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.