Fetching the paper…
Reading the bibliography…
Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems.
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 1908
Earlier work this paper cites.
Unpacking the Social Media Bot: A Typology to Guide Research and Policy
Robert Gorwa and Douglas Guilbeault · 1944
Earlier work this paper cites.
The intent to deceive
Roderick M Chisholm and Thomas D Feehan · 1977
Earlier work this paper cites.
Ignorance, Probability and Rational Choice on JSTOR
Isaac Levi · 1982
Earlier work this paper cites.
Decisions with indeterminate probabilities
Teddy Seidenfeld · 1983
Earlier work this paper cites.
Signaling games and stable equilibria
In-Koo Cho and David M Kreps · 1987
Earlier work this paper cites.
BDI Logics for BDI Architectures: Old Problems, New Perspectives
Andreas Herzig, Emiliano Lorini, Laurent Perrussel, and Zhanhao Xiao · 1987
Earlier work this paper cites.
Deception Games
V. J. Baston and F. A. Bostock · 1988
Earlier work this paper cites.
The peculiar effects of love and desire
Bas Van Fraassen · 1988
Earlier work this paper cites.
Intention is choice with commitment
Philip R. Cohen and Hector J. Levesque · 1990
Earlier work this paper cites.
Belief, Acceptance and Knowledge
L. Jonathan Cohen · 1995
Earlier work this paper cites.
The deceptive number changing game, in the absence of symmetry
Bert Fristedt · 1997
Earlier work this paper cites.
Intention
Gertrude Elizabeth Margaret Anscombe · 2000
Earlier work this paper cites.
Multi-agent influence diagrams for representing and solving games
Daphne Koller and Brian Milch · 2003
Earlier work this paper cites.
Learning in BDI Multi-agent Systems
Alejandro Guerra-Hernández et al · 2004
Earlier work this paper cites.
Learning Within the BDI Framework: An Empirical Analysis
Toan Phung et al · 2005
Earlier work this paper cites.
Evolutionary Conditions for the Emergence of Communication in Robots
Dario Floreano, Sara Mitri, Stéphane Magnenat, and Laurent Keller · 2007
Earlier work this paper cites.
On the reasoning patterns of agents in games
Avi Pfeffer and Ya’akov Gal · 2007
Earlier work this paper cites.
Epistemic logic and information update
Alexandru Baltag et al · 2008
Earlier work this paper cites.
Global catastrophic risks
Nick Bostrom Milan M Cirkovic · 2008
Earlier work this paper cites.
Ignorance and indifference*
John D. Norton · 2008
Earlier work this paper cites.
The basic AI drives
Stephen M. Omohundro · 2008
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
The concept of mind
Gilbert Ryle · 2009
Earlier work this paper cites.
Lying and deception: Theory and practice
Thomas L Carson · 2010
Earlier work this paper cites.
Open problems in cooperative ai, 2020
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R. McKee, Joel Z. Leibo, Kate Larson, and Thore Graepel · 2012
Earlier work this paper cites.
The Bayesian who knew too much
Yann Benétreau-Dupin · 2015
Earlier work this paper cites.
Automatic deception detection: Methods for finding fake news
Nadia K. Conroy, Victoria L. Rubin, and Yimin Chen · 2015
Earlier work this paper cites.
Inference of intention and permissibility in moral decision making
Max Kleiman-Weiner, Tobias Gerstenberg, Sydney Levine, and Joshua B. Tenenbaum · 2015
Earlier work this paper cites.
Hypergame theory: a model for conflict, misperception, and deception
Nicholas S Kovach et al · 2015
Earlier work this paper cites.
Deception in game theory: a survey and multiobjective model
Austin L Davis · 2016
Earlier work this paper cites.
Actual causality
Joseph Y Halpern · 2016
Earlier work this paper cites.
The Definition of Lying and Deception
James Edwin Mahon · 2016
Earlier work this paper cites.
Truth and probability
Frank P Ramsey · 2016
Cited alongside, same era.
Decisions and dependence in influence diagrams
Ross D. Shachter · 2016
Cited alongside, same era.
Machine learning with adversaries: Byzantine tolerant gradient descent
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer · 2017
Cited alongside, same era.
Deal or No Deal? End-to-End Learning for Negotiation Dialogues
Mike Lewis, Denis Yarats, Yann N. Dauphin, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Cited alongside, same era.
Artful paltering: The risks and rewards of using truthful statements to mislead others
Human-level play in the game of <i>diplomacy</i> by combining language models with strategic reasoning
Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra · 2022
Later among the works it cites.
Is power-seeking ai an existential risk?, 2022
Joseph Carlsmith · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Todd Rogers, Richard Zeckhauser, Francesca Gino, Michael I Norton, and Maurice E Schweitzer · 2017
Cited alongside, same era.
Certified defenses for data poisoning attacks
Jacob Steinhardt et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Zero-sum polymatrix games with link uncertainty: A Dempster-Shafer theory solution
Xinyang Deng et al · 2018
Cited alongside, same era.
What Ignorance Really Is. Examining the Foundations of Epistemology of Ignorance
Nadja El Kassar · 2018
Cited alongside, same era.
Towards formal definitions of blameworthiness, intention, and moral responsibility
Joseph Y. Halpern and Max Kleiman-Weiner · 2018
Cited alongside, same era.
Lies, bullshit, and deception in agent-oriented programming languages
Alison R. Panisson, Stefan Sarkadi, Peter McBurney, Simon Parsons, and Rafael H. Bordini · 2018
Cited alongside, same era.
Later among the works it cites.
ARC’s first technical report: Eliciting Latent Knowledge - AI Alignment Forum, May 2022
Paul Christiano · 2022
Later among the works it cites.
Path-Specific Objectives for Safer Agent Incentives
Sebastian Farquhar et al · 2022
Later among the works it cites.
E-Friend: A Logical-Based AI Agent System Chat-Bot for Emotional Well-Being and Mental Health
Mauricio J. Osorio Galindo, Luis A. Montiel Moreno, David Rojas-Velázquez, and Juan Carlos Nieves · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Later among the works it cites.
Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt · 2022
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin et al · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T. Connor, Neil Burch, Thomas Anthony, Stephen McAleer, Romuald Elie, Sarah H. Cen, Zhe Wang, Audrunas Gruslys, Aleksandra Malysheva, Mina Khan, Sherjil Ozair, Finbarr Timbers, Toby Pohlen, Tom Eccles, Mark Rowland, Marc Lanctot, Jean-Baptiste Lespiau, Bilal Piot, Shayegan Omidshafiei, Edward Lockhart, Laurent Sifre, Nathalie Beauguerlange, Remi Munos, David Silver, Satinder Singh, Demis Hassabis, and Karl Tuyls · 2022
Later among the works it cites.
Intention
Kieran Setiya · 2022
Later among the works it cites.
Talking about large language models, 2022
Murray Shanahan · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro · 2022
Later among the works it cites.
Manipulation and the ai act, 2022
Risto Uuk · 2022
Later among the works it cites.
On agent incentives to manipulate human feedback in multi-agent reward learning scenarios
Francis Rhys Ward et al · 2022
Later among the works it cites.
Frontier ai regulation: Managing emerging risks to public safety, 2023
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf · 2023
Closest in time.
Chatbots, deepfakes, and voice clones: AI deception for sale, March 2023
Michael Atleson · 2023
Closest in time.
Taken out of context: On measuring situational awareness in llms, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Closest in time.
Characterizing manipulation from ai systems, 2023
Micah Carroll, Alan Chan, Henry Ashton, and David Krueger · 2023
Closest in time.
Generative language models and automated influence operations: Emerging threats and potential mitigations, 2023
Josh A. Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Closest in time.
Context is environment, 2023
Sharut Gupta, Stefanie Jegelka, David Lopez-Paz, and Kartik Ahuja · 2023
Closest in time.
Reasoning about causality in games
Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, Alessandro Abate, and Michael Wooldridge · 2023
Closest in time.
An overview of catastrophic ai risks, 2023
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Closest in time.
Still no lie detector for language models: Probing empirical and conceptual roadblocks, 2023
B. A. Levinstein and Daniel A. Herrmann · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions, 2023
Lorenzo Pacchiardi, Alex J. Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y. Pan, Yarin Gal, Owain Evans, and Jan Brauner · 2023
Closest in time.
Ai deception: A survey of examples, risks, and potential solutions, 2023
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2023
Closest in time.
Experiments with detecting and mitigating ai deception, 2023
Ismail Sahbane, Francis Rhys Ward, and C Henrik Åslund · 2023
Closest in time.
Model evaluation for extreme risks, 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Closest in time.