Fetching the paper…
Reading the bibliography…
Recent capability increases in large language models (LLMs) open up applications in which groups of communicating generative AI agents solve joint tasks.
Hijacking Malaria Simulators with Probabilistic Programming, May 2019
Bradley Gram-Hansen, Christian Schroeder de Witt, Tom Rainforth, Philip H. S. Torr, Yee Whye Teh, and Atılım Güneş Baydin · 1905
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems, December 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Machine Unlearning, December 2020
Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 1912
Earlier work this paper cites.
An Unsolvable Problem of Elementary Number Theory
Alonzo Church · 1936
Earlier work this paper cites.
λ \lambda -definability and recursiveness
S. C. Kleene · 1936
Earlier work this paper cites.
On Computable Numbers, with an Application to the Entscheidungsproblem
A. M. Turing · 1937
Earlier work this paper cites.
The psychology of coordination and common knowledge
Kyle A. Thomas, Peter DeScioli, Omar Sultan Haque, and Steven Pinker · 1939
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
A Markovian Decision Process
Richard Bellman · 1957
Earlier work this paper cites.
Knowledge and Belief
Jaakko Hintikka · 1962
Earlier work this paper cites.
Secrecy, authentication, and public key systems
Ralph Charles Merkle · 1979
Earlier work this paper cites.
The computer as a physical system: A microscopic quantum mechanical Hamiltonian model of computers as represented by Turing machines
Paul Benioff · 1980
Earlier work this paper cites.
Relative to a Random Oracle A, ${\bf P}^A \ne {\bf NP}^A \ne \text{co-}{\bf NP}^A $ with Probability 1
Charles H. Bennett and John Gill · 1981
Earlier work this paper cites.
The Strategy of Conflict: With a New Preface by the Author
Thomas C. Schelling · 1981
Earlier work this paper cites.
The Prisoners’ Problem and the Subliminal Channel
Gustavus J. Simmons · 1984
Earlier work this paper cites.
How to generate and exchange secrets
Andrew Chi-Chih Yao · 1986
Earlier work this paper cites.
Using Reasoning About Knowledge to Analyze Distributed Systems
J Y Halpern · 1987
Earlier work this paper cites.
Training Stochastic Model Recognition Algorithms as Networks can Lead to Maximum Mutual Information Estimation of Parameters
John Bridle · 1989
Earlier work this paper cites.
Approximating common knowledge with common beliefs
Dov Monderer and Dov Samet · 1989
Earlier work this paper cites.
Learning How to Cooperate: Optimal Play in Repeated Coordination Games
Vincent P. Crawford and Hans Haller · 1990
Earlier work this paper cites.
Knowledge and common knowledge in a distributed environment
Joseph Y. Halpern and Yoram Moses · 1990
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnad · 1990
Earlier work this paper cites.
Probabilistic knowledge and probabilistic common knowledge
Paul Krasucki, Gilbert Ndjatou, and Rohit Parikh · 1990
Earlier work this paper cites.
Game Theory
Drew Fudenberg · 1991
Earlier work this paper cites.
Natural Language and Universal Grammar: Essays in Linguistic Theory , volume 1
John Lyons · 1991
Earlier work this paper cites.
The MD5 Message-Digest Algorithm
Ronald L. Rivest · 1992
Earlier work this paper cites.
Electronic Watermark
Andrew Tirkel, G.A. Rankin, Ron van Schyndel, W. Ho, N. Mee, and C. Osborne · 1994
Earlier work this paper cites.
Generative Quantum Machine Learning, November 2021
Christa Zoufal · 1994
Earlier work this paper cites.
An information-theoretic model for steganography
Christian Cachin · 1998
Earlier work this paper cites.
Interactive team reasoning: A contribution to the theory of co-operation
Michael Bacharach · 1999
Earlier work this paper cites.
Knowledge and common knowledge in a distributed environment, June 2000
Joseph Y. Halpern and Yoram Moses · 2000
Earlier work this paper cites.
Provably secure steganography
Nicholas J. Hopper, John Langford, and Luis von Ahn · 2002
Earlier work this paper cites.
OCB: A block-cipher mode of operation for efficient authenticated encryption
Phillip Rogaway, Mihir Bellare, and John Black · 2003
Earlier work this paper cites.
An information-theoretic model for steganography
Christian Cachin · 2004
Earlier work this paper cites.
Public-key steganography
Luis von Ahn and Nicholas J. Hopper · 2004
Earlier work this paper cites.
Public-Key Steganography
Luis von Ahn and Nicholas J. Hopper · 2004
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Introduction to modern cryptography : principles and protocols
Jonathan Katz and Yehuda Lindell · 2007
Earlier work this paper cites.
On the secure hash algorithm family
Wouter Penard and Tim van Werkhoven · 2008
Earlier work this paper cites.
Random oracles in a quantum world
Dan Boneh, Özgür Dagdelen, Marc Fischlin, Anja Lehmann, Christian Schaffner, and Mark Zhandry · 2011
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Variational Dropout and the Local Reparameterization Trick, December 2015
Diederik P. Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Low-Complexity Cryptographic Hash Functions
Benny Applebaum, Naama Haramaty-Krasne, Yuval Ishai, Eyal Kushilevitz, and Vinod Vaikuntanathan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Reframing superintelligence: Comprehensive ai services as general intelligence
K. Eric Drexler · 2019
Cited alongside, same era.
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, March 2019
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Cited alongside, same era.
Human Compatible
Stuart Russell · 2019
Cited alongside, same era.
Multi-Agent Common Knowledge Reinforcement Learning
Christian Schroeder de Witt, Jakob Foerster, Gregory Farquhar, Philip Torr, Wendelin Boehmer, and Shimon Whiteson · 2019
Cited alongside, same era.
Neural Linguistic Steganography
Zachary Ziegler, Yuntian Deng, and Alexander Rush · 2019
Undetectable Watermarks for Language Models, 2023
Miranda Christ, Sam Gunn, and Or Zamir · 2023
Later among the works it cites.
Emilio Ferrara · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini-Team · 2023
Later among the works it cites.
Ai control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2023
Later among the works it cites.
Intelligent Virtual Assistants with LLM-based Process Automation, December 2023
Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, Longfei Li, Jinjie Gu, and Chenyi Zhuang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Artificial Intelligence: Can Seemingly Collusive Outcomes Be Avoided?, February 2020
Ibrahim Abada and Xavier Lambin · 2020
Cited alongside, same era.
Communication complexity of Nash equilibrium in potential games (extended abstract)
Yakov Babichenko and Aviad Rubinstein · 2020
Cited alongside, same era.
Artificial Intelligence, Algorithmic Pricing, and Collusion
Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello · 2020
Cited alongside, same era.
Recent Advances of Image Steganography With Generative Adversarial Networks
Jia Liu, Yan Ke, Zhuo Zhang, Yu Lei, Jun Li, Minqing Zhang, and Xiaoyuan Yang · 2020
Cited alongside, same era.
Fast agent-based simulation framework with applications to reinforcement learning and the study of trading latency effects
Peter Belcak, Jan-Peter Calliess, and Stefan Zohren · 2021
Cited alongside, same era.
QCB: Efficient Quantum-secure Authenticated Encryption, 2020
Ritam Bhaumik, Xavier Bonnetain, André Chailloux, Gaëtan Leurent, María Naya-Plasencia, André Schrottenloher, and Yannick Seurin · 2021
Cited alongside, same era.
Mohsen Jamali, Ziv M. Williams, and Jing Cai · 2023
Later among the works it cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
On the Reliability of Watermarks for Large Language Models, June 2023
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein · 2023
Later among the works it cites.
Theory of Mind Might Have Spontaneously Emerged in Large Language Models, November 2023
Michal Kosinski · 2023
Later among the works it cites.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Later among the works it cites.
Camel: Communicative agents for" mind" exploration of large scale language model society
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Later among the works it cites.
Concise and Efficient Quantum Algorithms for Distribution Closeness Testing, February 2023
Lvzhou Li and Jingquan Luo · 2023
Later among the works it cites.
Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement Learning, March 2023
Yat Long Lo, Christian Schroeder de Witt, Samuel Sokota, Jakob Nicolaus Foerster, and Shimon Whiteson · 2023
Later among the works it cites.
Tell, don’t show: Declarative facts influence how llms generalize
Alexander Meinke and Owain Evans · 2023
Later among the works it cites.
A perfect collusion benchmark: How can ai agents be prevented from colluding with information-theoretic undetectability?
Sumeet Ramesh Motwani, Mikhail Baranchuk, Lewis Hammond, and Christian Schroeder de Witt · 2023
Later among the works it cites.
Privacy Issues in Large Language Models: A Survey, December 2023
Seth Neel and Peter Chang · 2023
Later among the works it cites.
Open X-Embodiment: Robotic Learning Datasets and RT-X Models, December 2023
Open X-Embodiment Collaboration, Abhishek Padalkar, Acorn Pooley, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Animesh Garg, Anthony Brohan, Antonin Raffin, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Brian Ichter, and Cewu Lu · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
GPT-4 Technical Report, December 2023
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, and Jake Berdine · 2023
Later among the works it cites.
Language Model Tokenizers Introduce Unfairness Between Languages, October 2023
Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, and Adel Bibi · 2023
Later among the works it cites.
Communicative Agents for Software Development, December 2023
Chen Qian, Xin Cong, Wei Liu, Cheng Yang, Weize Chen, Yusheng Su, Yufan Dang, Jiahao Li, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs, October 2023
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Preventing Language Models From Hiding Their Reasoning
Fabien Roger and Ryan Greenblatt · 2023
Later among the works it cites.
Are Emergent Abilities of Large Language Models a Mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Later among the works it cites.
Multi-Agent Security Workshop at NeurIPS 2023, 2023a
Christian Schroeder de Witt, Hawra Milani, Klaudia Krawiecka, Swapneel Mehta, Carla Cremer, and Martin Strohmeier · 2023
Later among the works it cites.
Cooperative AI via Decentralized Commitment Devices, November 2023
Xinyuan Sun, Davide Crapis, Matt Stephenson, Barnabé Monnot, Thomas Thiery, and Jonathan Passerat-Palmbach · 2023
Later among the works it cites.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, October 2023
The White House · 2023
Later among the works it cites.
Alexander Sasha Vezhnevets, John P. Agapiou, Avia Aharon, Ron Ziv, Jayd Matyas, Edgar A. Duéñez-Guzmán, William A. Cunningham, Simon Osindero, Danny Karmon, and Joel Z. Leibo · 2023
Later among the works it cites.
Honesty Is the Best Policy: Defining and Mitigating AI Deception, December 2023
Francis Rhys Ward, Francesco Belardinelli, Francesca Toni, and Tom Everitt · 2023
Later among the works it cites.
Mindstorms in Natural Language-Based Societies of Mind, May 2023
Mingchen Zhuge, Haozhe Liu, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Abed Al Kader Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li, Guohao Li, Shuming Liu, Jinjie Mai, Piotr Piękos, Aditya Ramesh, Imanol Schlag, Weimin Shi, Aleksandar Stanić, Wenyi Wang, Yuhui Wang, Mengmeng Xu, Deng-Ping Fan, Bernard Ghanem, and Jürgen Schmidhuber · 2023
Later among the works it cites.
Foundational Challenges in Assuring Alignment and Safety of Large Language Models, April 2024
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Edelman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric Bigelow, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Zhang, Ruiqi Zhong, Seán Ó hÉigeartaigh, Gabriel Recchia, Giulio Corsi, Alan Chan, Markus Anderljung, Lilian Edwards, Yoshua Bengio, Danqi Chen, Samuel Albanie, Tegan Maharaj, Jakob Foerster, Florian Tramer, He He, Atoosa Kasirzadeh, Yejin Choi, and David Krueger · 2024
Closest in time.
Unelicitable backdoors in language models via cryptographic transformer circuits, 2024
Andis Draguns, Andrew Gritsevskiy, Sumeet Ramesh Motwani, Charlie Rogers-Smith, Jeffrey Ladish, and Christian Schroeder de Witt · 2024
Closest in time.
Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks, May 2024
Tim Franzmeyer, Stephen McAleer, João F. Henriques, Jakob N. Foerster, Philip H. S. Torr, Adel Bibi, and Christian Schroeder de Witt · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, January 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
A gentle introduction to mechanistic anomaly detection, April 2024
Erik Jenner · 2024
Closest in time.
Evaluating language-model agents on realistic autonomous tasks, 2024
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R. Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, and Paul Christiano · 2024
Closest in time.
Eliciting latent knowledge from quirky language models, 2024
Alex Mallen, Madeline Brumley, Julia Kharchenko, and Nora Belrose · 2024
Closest in time.
Linas Nasvytis, Kai Sandbrink, Jakob Foerster, Tim Franzmeyer, and Christian Schroeder de Witt · 2024
Closest in time.
Introducing the GPT Store, January 2024
OpenAI · 2024
Closest in time.
AI deception: A survey of examples, risks, and potential solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks · 2024
Closest in time.
Prompting a pretrained transformer can be a universal approximator, 2024
Aleksandar Petrov, Philip H. S. Torr, and Adel Bibi · 2024
Closest in time.
Can Large Language Model Agents Simulate Human Trust Behaviors?, March 2024
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li · 2024
Closest in time.