Fetching the paper…
Reading the bibliography…
This paper argues that a range of current AI systems have learned how to deceive humans.
“Fine-tuning language models from human preferences”, 2020
Daniel. Ziegler et al · 1909
Earlier work this paper cites.
“How to define theoretical terms”
David Lewis · 1970
Earlier work this paper cites.
“Inquiry”
Robert Stalnaker · 1984
Earlier work this paper cites.
“Psychosemantics: The problem of meaning in the philosophy of mind”
Jerry. Fodor · 1987
Earlier work this paper cites.
“Influence tactics, affect, and exchange quality in supervisor-subordinate interactions: A laboratory experiment and field study.”
Sandy Wayne and Gerald Ferris · 1990
Earlier work this paper cites.
“Impact of ingratiation on judgments and evaluations: A meta-analytic investigation.”
Randall Gordon · 1996
Earlier work this paper cites.
“Deception in war”
Jon Latimer · 2001
Earlier work this paper cites.
“Bald-faced lies! Lying without the intent to deceive”
Roy Sorensen · 2007
Earlier work this paper cites.
“The basic AI drives”
Stephen. Omohundro · 2008
Earlier work this paper cites.
“Deceit and self-deception: Fooling yourself the better to fool others”
Robert Trivers · 2011
Earlier work this paper cites.
“Is US economic growth over? Faltering innovation confronts the six headwinds”, 2012
Robert Gordon · 2012
Earlier work this paper cites.
“Ontological commitment”
Phillip Bricker · 2016
Earlier work this paper cites.
“The definition of lying and deception”
James Mahon · 2016
Earlier work this paper cites.
Paul Christiano et al · 2017
Earlier work this paper cites.
“Deal or no deal? End-to-end learning for negotiation dialogues”, 2017
Mike Lewis et al · 2017
Earlier work this paper cites.
“OpenAI charter”
OpenAI · 2018
Earlier work this paper cites.
“Superhuman AI for multiplayer poker”
Noam Brown and Tuomas Sandholm · 2019
Earlier work this paper cites.
“Carnegie Mellon and Facebook AI beats professionals in six-player poker” Accessed: 27 July 2023, 2019
Carnegie Mellon University · 2019
Earlier work this paper cites.
“Toronto van attack suspect says he was ‘radicalized’ online by ‘incels”’
Leyland Cecco · 2019
Earlier work this paper cites.
“The Volkswagen emissions scandal and its aftermath”
Jae. Jung and Elizabeth Sharon · 2019
Earlier work this paper cites.
“StarCraft is a deep, complicated war strategy game. Google’s AlphaStar AI crushed it.”
Kelsey Piper · 2019
Earlier work this paper cites.
“Human compatible: Artificial intelligence and the problem of control”
Stuart Russell · 2019
Earlier work this paper cites.
“Fraudsters used AI to mimic CEO’s voice in unusual cybercrime case”
Catherine Stupp · 2019
Earlier work this paper cites.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning”
Oriol Vinyals et al · 2019
Earlier work this paper cites.
“Multiple realizability”
John Bickle · 2020
Earlier work this paper cites.
“The alignment problem: Machine learning and human values”
Brian Christian · 2020
Earlier work this paper cites.
“Aligning AI with shared human values”
Dan Hendrycks et al · 2020
Earlier work this paper cites.
“The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life research communities”
Joel Lehman et al · 2020
Earlier work this paper cites.
“Transformers: State-of-the-art natural language processing”
Thomas Wolf et al · 2020
Earlier work this paper cites.
“A general language assistant as a laboratory for alignment”, 2021
Amanda Askell et al · 2021
Earlier work this paper cites.
“On the dangers of stochastic parrots: Can language models be too big?”
Emily. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
“Proposal for a regulation of the European Parliament and of the council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) and amending certain union legislative acts” COM(2021) 206 final, 2021/0106 (COD), SEC(2021) 167 final; SWD(2021) 84 final; SWD(2021) 85 final, 2021
European Commission · 2021
Earlier work this paper cites.
“Truthful AI: Developing and governing AI that does not lie”, 2021
Owain Evans et al · 2021
Earlier work this paper cites.
“Artificial Intelligence Act”, 2023
Tambiama Madiega · 2021
Earlier work this paper cites.
“Characterising deception in AI: A survey”
Peta Masters, Wally Smith, Liz Sonenberg and Michael Kirley · 2021
Earlier work this paper cites.
“Constitutional AI: Harmlessness from AI feedback”, 2022
Yuntao Bai et al · 2022
Cited alongside, same era.
“Game 438141. Cicero is FRANCE. Dialogue with E,G,R shown.”, Data relevant to Bakhtin et al., 2022b., 2022
Anton Bakhtin et al · 2022
Cited alongside, same era.
“Human-level play in the game of Diplomacy by combining language models with strategic reasoning”
Anton Bakhtin et al · 2022
Cited alongside, same era.
“Cicero playing as Austria sure seems like they manipulated/decieved a human Russia and are now justifying it [Tweet]”, Twitter, 2022
Haydn Belfield · 2022
Cited alongside, same era.
“Discovering latent knowledge in language models without supervision”, 2022
Collin Burns, Haotian Ye, Dan Klein and Jacob Steinhardt · 2022
Cited alongside, same era.
“Evaluating language-model agents on realistic autonomous tasks”, 2023
Megan Kinniment et al · 2023
Closest in time.
“A watermark for large language models”, 2023
John Kirchenbauer et al · 2023
Closest in time.
“Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense”, 2023
Kalpesh Krishna et al · 2023
Closest in time.
“Werewolf among us: Multimodal resources for modeling persuasion behaviors in social deduction games”
Bolin Lai et al · 2023
Closest in time.
“Goal misgeneralization in deep reinforcement learning”, 2023
Lauro Langosco et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Reality+: Virtual worlds and the problems of philosophy”
David Chalmers · 2022
Cited alongside, same era.
“Our infra went down for 10 minutes and Cicero (France) explains its absence (lol) [Tweet]”, Twitter, 2022
Emily Dinan · 2022
Cited alongside, same era.
“It’s designed to never intentionally backstab - all its messages correspond to actions it currently plans to take. [Tweet]”, Twitter, 2022
Mike Lewis · 2022
Cited alongside, same era.
“TruthfulQA: Measuring how models mimic human falsehoods”, 2022
Stephanie Lin, Jacob Hilton and Owain Evans · 2022
Cited alongside, same era.
“Introducing ChatGPT”, 2022
OpenAI · 2022
Cited alongside, same era.
“Discovering language model behaviors with model-written evaluations”, 2022
Ethan Perez et al · 2022
Cited alongside, same era.
“Goal misgeneralization: Why correct specifications aren’t enough for correct goals”, 2022
Rohin Shah et al · 2022
Cited alongside, same era.
“Measuring faithfulness in chain-of-thought reasoning”
Tamera Lanham et al · 2023
Closest in time.
“Functionalism”
Janet Levin · 2023
Closest in time.
“Still no lie detector for language Models: Probing empirical and conceptual roadblocks”, 2023
B.. Levinstein and Daniel. Herrmann · 2023
Closest in time.
“Inference-time intervention: Eliciting truthful answers from a language model”, 2023
Kenneth Li et al · 2023
Closest in time.
“AgentBench: Evaluating LLMs as agents”, 2023
Xiao Liu et al · 2023
Closest in time.
“Did GPT-4 hire and then lie to a task rabbit worker to solve a CAPTCHA?”
Melanie Mitchell · 2023
Closest in time.
“That story about a killer AI run amok seems fake. [Tweet]”, Twitter, 2023
Dan Neidle · 2023
Closest in time.
“Hoodwinked: Deception and cooperation in a text-based game for language models”, 2023
Aidan O’Gara · 2023
Closest in time.
“GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses”, 2023
OpenAI · 2023
Closest in time.
“GPT-4 technical report”, 2023
OpenAI · 2023
Closest in time.
Alexander Pan et al · 2023
Closest in time.
“How AI puts elections at risk — and the needed safeguards”
Mekela Panditharatne and Noah Giansiracusa · 2023
Closest in time.
“Can AI-generated text be reliably detected?”, 2023
Vinu Sadasivan et al · 2023
Closest in time.
“Evaluating the moral beliefs encoded in LLMs”, 2023
Nino Scherrer, Claudia Shi, Amir Feder and David. Blei · 2023
Closest in time.
“Emergent deception and skepticism via theory of mind”
Lion Schulz, Nitay Alon, Jeffrey Rosenschein and Peter Dayan · 2023
Closest in time.
“Role-play with large language models”, 2023
Murray Shanahan, Kyle McDonell and Laria Reynolds · 2023
Closest in time.
“The Gaslighting Among Us AI [Video]”, 2023
Tim Shaw · 2023
Closest in time.
“Model evaluation for extreme risks”, 2023
Toby Shevlane et al · 2023
Closest in time.
“Playing the Werewolf game with artificial intelligence for language understanding”, 2023
Hisaichi Shibata, Soichiro Miki and Yuta Nakamura · 2023
Closest in time.
“Can large language models democratize access to dual-use biotechnology?”, 2023
Emily. Soice et al · 2023
Closest in time.
“Emergent deception and emergent optimization”, 2023
J Steinhardt · 2023
Closest in time.
“‘A relationship with another human is overrated’ – inside the rise of AI girlfriends”
James Titcomb · 2023
Closest in time.
“AI poses national security threat, warns terror watchdog”
Mark Townsend · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez and Samuel. Bowman · 2023
Closest in time.
“They thought loved ones were calling for help. It was an AI scam.”
Pranshu Verma · 2023
Closest in time.
“A.I. is helping hackers make better phishing emails”
Bob Violino · 2023
Closest in time.
“Startup uses AI chatbot to provide mental health counseling and then realizes it ‘feels weird”’
Chloe Xiang · 2023
Closest in time.
“Representation engineering: Understanding and controlling the inner workings of neural networks” Manuscript, 2023
Andy Zou et al · 2023
Closest in time.
“Universal and transferable adversarial attacks on aligned language models”, 2023
Andy Zou, Zifan Wang, J. Kolter and Matt Fredrikson · 2023
Closest in time.