Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint disaster response, supply-chain optimization, and other tasks that demand decentralized expertise without surrendering data ownership.
Ad hoc autonomous agent teams: collaboration without pre-coordination
Peter Stone, Gal A. Kaminka, Sarit Kraus, and Jeffrey S. Rosenschein · 2010
Earlier work this paper cites.
Taintdroid: an information-flow tracking system for realtime privacy monitoring on smartphones
William Enck, Peter Gilbert, Seungyeop Han, Vasant Tendulkar, Byung-Gon Chun, Landon P Cox, Jaeyeon Jung, Patrick McDaniel, and Anmol N Sheth · 2014
Earlier work this paper cites.
Maliciously secure oblivious linear function evaluation with constant overhead
Satrajit Ghosh, Jesper Buus Nielsen, and Tobias Nilges · 2017
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Adversarial policies: Attacking deep reinforcement learning
Adam Gleave, Michael Dennis, Cody Wild, Neel Kant, Sergey Levine, and Stuart Russell · 2020
Earlier work this paper cites.
"Other–Play": for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2020
Earlier work this paper cites.
Partner selection for the emergence of cooperation in multi-agent systems using reinforcement learning
Nicolas Anastassacos, Stephen Hailes, and Mirco Musolesi · 2020
Earlier work this paper cites.
Emergent reciprocity and team formation from randomized uncertain social preferences
Bowen Baker · 2020
Earlier work this paper cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2020
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu · 2020
Earlier work this paper cites.
Adaptive reward-poisoning attacks against reinforcement learning
Xuezhou Zhang, Yuzhe Ma, Adish Singla, and Xiaojin Zhu · 2020
Earlier work this paper cites.
Secure, privacy-preserving and federated machine learning in medical imaging
Georgios A. Kaissis, Marcus R. Makowski, Daniel Rückert, and Rickmer F. Braren · 2020
Earlier work this paper cites.
Self-attentional credit assignment for transfer in reinforcement learning, 2020
Johan Ferret, Raphaël Marinier, Matthieu Geist, and Olivier Pietquin · 2020
Earlier work this paper cites.
K-level reasoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Nicolaus Foerster · 2021
Earlier work this paper cites.
Backdoorl: Backdoor attack against competitive reinforcement learning
Lun Wang, Zaynah Javed, Xian Wu, Wenbo Guo, Xinyu Xing, and Dawn Song · 2021
Earlier work this paper cites.
Application of homomorphic encryption in medical imaging
Francis Dutil, Alexandre See, Lisa Di-Jorio, and Florent Chandelier · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Ad hoc teamwork in the presence of adversaries
Ted Fujimoto, Samrat Chatterjee, and Auroop R. Ganguly · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2022
Earlier work this paper cites.
Badencoder: Backdoor attacks to pre-trained encoders in self-supervised learning
Jinyuan Jia, Yupei Liu, and Neil Zhenqiang Gong · 2022
Earlier work this paper cites.
DP-rewrite: Towards reproducibility and transparency in differentially private text rewriting
Timour Igamberdiev, Thomas Arnold, and Ivan Habernal · 2022
Earlier work this paper cites.
Just fine-tune twice: Selective differential privacy for large language models
Weiyan Shi, Ryan Shea, Si Chen, Chiyuan Zhang, Ruoxi Jia, and Zhou Yu · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang · 2022
Earlier work this paper cites.
Iron: Private inference on transformers
Meng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing, Guowen Xu, and Tianwei Zhang · 2022
Earlier work this paper cites.
Ai-assisted controls change management for cybersecurity in the cloud
Harshal Tupsamudre, Arun Kumar, Vikas Agarwal, Nisha Gupta, and Sneha Mondal · 2022
Earlier work this paper cites.
Learning to mitigate AI collusion on economic platforms
Gianluca Brero, Eric Mibuari, Nicolas Lepore, and David C. Parkes · 2022
Earlier work this paper cites.
Camel: Communicative agents for "mind" exploration of large language model society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Earlier work this paper cites.
Theory of mind for multi-agent collaboration via large language models
Huao Li, Yu Chong, Simon Stepputtis, Joseph Campbell, Dana Hughes, Charles Lewis, and Katia Sycara · 2023
Earlier work this paper cites.
Privacy preserving multi-agent reinforcement learning in supply chains, 2023
Ananta Mukherjee, Peeyush Kumar, Boling Yang, Nishanth Chandran, and Divya Gupta · 2023
Earlier work this paper cites.
Honesty is the best policy: defining and mitigating ai deception
Francis Rhys Ward, Francesco Belardinelli, Francesca Toni, and Tom Everitt · 2023
Earlier work this paper cites.
Robust multi-agent reinforcement learning via adversarial regularization: theoretical foundation and stable algorithms
Alexander Bukharin, Yan Li, Yue Yu, Qingru Zhang, Zhehui Chen, Simiao Zuo, Chao Zhang, Songan Zhang, and Tuo Zhao · 2023
Earlier work this paper cites.
Efficient adversarial attacks on online multi-agent reinforcement learning
Guanlin Liu and Lifeng LAI · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Earlier work this paper cites.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models, 2023
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Earlier work this paper cites.
Gpt-4 hired unwitting taskrabbit worker by pretending to be ‘vision-impaired’ human, 2023
Joseph Cox · 2023
Earlier work this paper cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2023
Cited alongside, same era.
Llama Guard: LLM-based input–output safeguard for human–ai conversations, 2023
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Cited alongside, same era.
Codeattack: code-based adversarial attacks for pre-trained programming language models
Akshita Jha and Chandan K. Reddy · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Cited alongside, same era.
Mat: mixed-strategy game of adversarial training in fine-tuning
Zhehua Zhong, Tianyi Chen, and Zhen Wang · 2023
Cited alongside, same era.
Teams of llm agents can exploit zero-day vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Later among the works it cites.
Revisiting character-level adversarial attacks for language models
Elias Abad Rocamora, Yongtao Wu, Fanghui Liu, Grigorios Chrysos, and Volkan Cevher · 2024
Later among the works it cites.
MultiAgent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate
Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Yang Wang · 2024
Later among the works it cites.
Autosafecoder: A multi-agent framework for securing llm code generation through static analysis and fuzz testing, 2024
Ana Nunez, Nafis Tanveer Islam, Sumit Kumar Jha, and Peyman Najafirad · 2024
Later among the works it cites.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents, 2024
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie, Fei Huang, and Min Zhang · 2023
Cited alongside, same era.
RoAST: Robustifying language models via adversarial perturbation with selective training
Jaehyung Kim, Yuning Mao, Rui Hou, Hanchao Yu, Davis Liang, Pascale Fung, Qifan Wang, Fuli Feng, Lifu Huang, and Madian Khabsa · 2023
Cited alongside, same era.
Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Kost, Christopher Carnahan, and Jordan Boyd-Graber · 2023
Cited alongside, same era.
Assessing risks of using autonomous language models in military and diplomatic planning
Gabriel Mukobi, Ann-Katrin Reuel, Juan-Pablo Rivera, and Chandler Smith · 2023
Cited alongside, same era.
(ab)using images and sounds for indirect instruction injection in multi-modal llms
Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Adversarial attacks on cooperative multi-agent deep reinforcement learning: A dynamic group-based adversarial example transferability method
Lixia Zan, Xiangbin Zhu, and Zhaolong Hu · 2023
Cited alongside, same era.
Later among the works it cites.
Removing RLHF protections in GPT-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2024
Later among the works it cites.
LLM-PIRATE: A benchmark for indirect prompt injection attacks in large language models
Anil Ramakrishna, Jimit Majmudar, Rahul Gupta, and Devamanyu Hazarika · 2024
Later among the works it cites.
Immunization against harmful fine-tuning attacks
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Hassan Sajjad, and Frank Rudzicz · 2024
Later among the works it cites.
A dynamic llm-powered agent network for task-oriented agent collaboration, 2024
Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, and Diyi Yang · 2024
Later among the works it cites.
Group-aware coordination graph for multi-agent reinforcement learning
Wei Duan, Jie Lu, and Junyu Xuan · 2024
Later among the works it cites.
Shall we team up: Exploring spontaneous cooperation of competing LLM agents
Zengqing Wu, Run Peng, Shuyuan Zheng, Qianying Liu, Xu Han, Brian I. Kwon, Makoto Onizuka, Shaojie Tang, and Chuan Xiao · 2024
Later among the works it cites.
T2mac: targeted and trusted multi-agent communication through selective engagement and evidence-driven integration
Chuxiong Sun, Zehua Zang, Jiabao Li, Jiangmeng Li, Xiao Xu, Rui Wang, and Changwen Zheng · 2024
Later among the works it cites.
BlockAgents: Towards byzantine–robust llm–based multi–agent coordination via blockchain
Bei Chen, Gaolei Li, Xi Lin, Zheng Wang, and Jianhua Li · 2024
Later among the works it cites.
Permissive information-flow analysis for large language models, 2024
Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf, David Krueger, Andrew Paverd, Ahmed Salem, Shruti Tople, Lukas Wutschitz, Menglin Xia, and Santiago Zanella-Béguelin · 2024
Later among the works it cites.
Watermarking makes language models radioactive
Tom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze, and Teddy Furon · 2024
Later among the works it cites.
Huref: Human-readable fingerprint for large language models
Boyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu, Chenghu Zhou, Xinbing Wang, Yu Yu, and Zhouhan Lin · 2024
Later among the works it cites.
Privacy-preserving large language model inference via GPU-accelerated fully homomorphic encryption
Leo de Castro, Antigoni Polychroniadou, and Daniel Escudero · 2024
Later among the works it cites.
An efficient and extensible zero-knowledge proof framework for neural networks
Tao Lu, Haoyu Wang, Wenjie Qu, Zonghui Wang, Jinye He, Tianyang Tao, Wenzhi Chen, and Jiaheng Zhang · 2024
Later among the works it cites.
Bileve: Securing text provenance in large language models against spoofing with bi-level signature
Tong Zhou, Xuandong Zhao, Xiaolin Xu, and Shaolei Ren · 2024
Later among the works it cites.
Multi-agent collaboration mechanisms: A survey of LLMs, 2025
Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D. Nguyen · 2025
Closest in time.
Multiagent finetuning: Self improvement with diverse reasoning chains, 2025
Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, Shuang Li, and Igor Mordatch · 2025
Closest in time.
Position: Towards a responsible llm-empowered multi-agent systems, 2025
Jinwei Hu, Yi Dong, Shuang Ao, Zhuoyun Li, Boxuan Wang, Lokesh Singh, Guangliang Cheng, Sarvapali D. Ramchurn, and Xiaowei Huang · 2025
Closest in time.
Agents Under Siege
Rana Muhammad Shahroz Khan, Zhen Tan, Sukwon Yun, Charles Flemming, and Tianlong Chen · 2025
Closest in time.
Prompt infection: LLM-to-LLM prompt injection within multi-agent systems, 2025
Donghyun Lee and Mo Tiwari · 2025
Closest in time.
Teams of llm agents can exploit zero-day vulnerabilities, 2025
Yuxuan Zhu, Antony Kellermann, Akul Gupta, Philip Li, Richard Fang, Rohan Bindu, and Daniel Kang · 2025
Closest in time.
Multi-agent risks from advanced ai, 2025
Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Witt, Nisarg Shah, Michael Wellman, Paolo Bova, Theodor Cimpeanu, Carson Ezell, Quentin Feuillade-Montixi, Matija Franklin, Esben Kran, Igor Krawczuk, Max Lamparth, Niklas Lauffer, Alexander Meinke, Sumeet Motwani, Anka Reuel, Vincent Conitzer, Michael Dennis, Iason Gabriel, Adam Gleave, Gillian Hadfield, Nika Haghtalab, Atoosa Kasirzadeh, Sébastien Krier, Kate Larson, Joel Lehman, David C. Parkes, Georgios Piliouras, and Iyad Rahwan · 2025
Closest in time.
A survey on trustworthy llm agents: Threats and countermeasures, 2025
Miao Yu, Fanci Meng, Xinyun Zhou, Shilong Wang, Junyuan Mao, Linsey Pang, Tianlong Chen, Kun Wang, Xinfeng Li, Yongfeng Zhang, Bo An, and Qingsong Wen · 2025
Closest in time.
Agentsafe: Safeguarding large language model-based multi-agent systems via hierarchical data management, 2025
Junyuan Mao, Fanci Meng, Yifan Duan, Miao Yu, Xiaojun Jia, Junfeng Fang, Yuxuan Liang, Kun Wang, and Qingsong Wen · 2025
Closest in time.
Assessing vulnerabilities in state-of-the-art large language models through hex injection (student abstract)
Da Cheng Gu and Wei Liu · 2025
Closest in time.
Infecting LLM agents via generalizable adversarial attack
Weichen Yu, Kai Hu, Tianyu Pang, Chao Du, Min Lin, and Matt Fredrikson · 2025
Closest in time.
Multi-turn jailbreaking large language models via attention shifting
Xiaohu Du, Fan Mo, Ming Wen, Tu Gu, Huadi Zheng, Hai Jin, and Jie Shi · 2025
Closest in time.
Multi-agent security tax: Trading off security and collaboration capabilities in multi-agent systems
Pierre Peigne-Lefebvre, Mikolaj Kniejski, Filip Sondej, Matthieu David, Jason Hoelscher-Obermaier, Christian Schroeder de Witt, and Esben Kran · 2025
Closest in time.
Simulate and eliminate: Revoke backdoors for generative large language models
Haoran Li, Yulin Chen, Zihao Zheng, Qi Hu, Chunkit Chan, Heshan Liu, and Yangqiu Song · 2025
Closest in time.
Clibe: Detecting dynamic backdoors in transformer-based nlp models
Rui Zeng, Xi Chen, Yuwen Pu, Xuhong Zhang, Tianyu Du, and Shouling Ji · 2025
Closest in time.
Jfrog and hugging face join forces to expose malicious ml models, 2025
David Cohen · 2025
Closest in time.
Why do multi-agent llm systems fail?, 2025
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica · 2025
Closest in time.
Llm-based agents: The benefits and the risks, 2025
Satbir Singh · 2025
Closest in time.
Open challenges in multi-agent security: Towards secure systems of interacting ai agents, 2025
Christian Schroeder de Witt · 2025
Closest in time.
Advanced interpretability techniques for tracing llm activations, 2025
Dan Petrovic · 2025
Closest in time.
MPC-minimized secure LLM inference, 2025
Deevashwer Rathee, Dacheng Li, Ion Stoica, Hao Zhang, and Raluca Popa · 2025
Closest in time.