Fetching the paper…
Reading the bibliography…
A high volume of recent ML security literature focuses on attacks against aligned large language models (LLMs).
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Data poisoning attacks against federated learning systems
Vale Tolpegin, Stacey Truex, Mehmet Emre Gursoy, and Ling Liu · 2020
Earlier work this paper cites.
Just how toxic is data poisoning? a unified benchmark for backdoor and data poisoning attacks
Avi Schwarzschild, Micah Goldblum, Arjun Gupta, John P Dickerson, and Tom Goldstein · 2021
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu · 2022
Earlier work this paper cites.
Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses
Micah Goldblum, Dimitris Tsipras, Chulin Xie, Xinyun Chen, Avi Schwarzschild, Dawn Song, Aleksander Madry, Bo Li, and Tom Goldstein · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Andres M. Bran, Sam Cox, Oliver Schilter, Andrew D. White, and Philippe Schwaller · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust · 2023
Earlier work this paper cites.
Catastrophic jailbreak of open-source llms via exploiting generation
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Earlier work this paper cites.
Jailbreaking chatgpt via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, Kailong Wang, and Yang Liu · 2023
Earlier work this paper cites.
Scalable extraction of training data from (production) language models
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A Feder Cooper, Daphne Ippolito, Christopher A Choquette-Choo, Eric Wallace, Florian Tramèr, and Katherine Lee · 2023
Earlier work this paper cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks, 2023
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh · 2023
Earlier work this paper cites.
Evil geniuses: Delving into the safety of llm-based agents
Yu Tian, Xiao Yang, Jingyuan Zhang, Yinpeng Dong, and Hang Su · 2023
Cited alongside, same era.
Llm-powered autonomous agents
Lilian Weng · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku
Anthropic · 2024
Cited alongside, same era.
Agent q: Advanced reasoning and learning for autonomous ai agents
Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov · 2024
Later among the works it cites.
Ransomware attack has cost unitedhealth $872 million; total expected to surpass $1 billion, 2024
The Record · 2024
Later among the works it cites.
“ do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Later among the works it cites.
Language agents achieve superhuman synthesis of scientific knowledge
Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White · 2024
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Breaking down the defenses: A comparative survey of attacks on large language models, 2024
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vaibhav Kumar, Vinija Jain, and Aman Chadha · 2024
Cited alongside, same era.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al · 2024
Cited alongside, same era.
Security and privacy challenges of large language models: A survey, 2024
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu · 2024
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents, 2024
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Cited alongside, same era.
Project mariner
Google DeepMind · 2024
Cited alongside, same era.
Urgent security alert for fedora linux 40 and fedora rawhide users, 2024
Red Hat · 2024
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Cited alongside, same era.
Later among the works it cites.
Trojllm: A black-box trojan prompt attack on large language models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou · 2024
Later among the works it cites.
Watch out for your agents! investigating backdoor threats to llm-based agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun · 2024
Later among the works it cites.
The good and the bad: Exploring privacy issues in retrieval-augmented generation (rag)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, et al · 2024
Later among the works it cites.
Breaking agents: Compromising autonomous llm agents through malfunction amplification
Boyang Zhang, Yicong Tan, Yun Shen, Ahmed Salem, Michael Backes, Savvas Zannettou, and Yang Zhang · 2024
Later among the works it cites.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Later among the works it cites.
We are now please: A new consumer ai company
MultiOn · 2025
Closest in time.
Introducing operator
OpenAI · 2025
Closest in time.