Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have gained widespread adoption across diverse applications due to their impressive generative capabilities.
Model inversion attacks that exploit confidence information and basic countermeasures
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart · 2015
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team and Google · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Earlier work this paper cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh · 2023
Earlier work this paper cites.
Data poisoning in LLMs: Jailbreak-tuning and scaling laws
Dillon Bowen, Brendan Murphy, Will Cai, David Khachaturov, Adam Gleave, and Kellin Pelrine · 2024
Cited alongside, same era.
Poisonbench: Assessing large language model vulnerability to data poisoning
Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B Cohen, David Krueger, and Fazl Barez · 2024
Cited alongside, same era.
On the (in) security of LLM app stores
Xinyi Hou, Yanjie Zhao, and Haoyu Wang · 2024
Cited alongside, same era.
Strengthening LLM ecosystem security: Preventing mobile malware from manipulating llm-based applications
Lu Huang, Jingfeng Xue, Yong Wang, Junbao Chen, and Tianwei Lei · 2024
Cited alongside, same era.
LLM platform security: applying a systematic evaluation framework to OpenAI’s chatgpt plugins
Adobe firefly
Adobe · 2025
Closest in time.
Claude opus 4.1
Anthropic · 2025
Closest in time.
Hugging face – the ai community building the future
Hugging Face · 2025
Closest in time.
LangChain
Harrison Chase · 2025
Closest in time.
Promptflow
Microsoft · 2025
Closest in time.
Gpt-5 system card
OpenAI · 2025
Closest in time.
Benchmarking and defending against indirect prompt injection attacks on large language models
Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Umar Iqbal, Tadayoshi Kohno, and Franziska Roesner · 2024
Cited alongside, same era.
Fath: Authentication-based test-time defense against indirect prompt injection attacks
Jiongxiao Wang, Fangzhou Wu, Wendi Li, Jinsheng Pan, Edward Suh, Z Morley Mao, Muhao Chen, and Chaowei Xiao · 2024
Cited alongside, same era.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang · 2024
Cited alongside, same era.
Poisonedrag: Knowledge corruption attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson
Cited in the paper.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson
Cited in the paper.
Qiusi Zhan, Richard Fang, Henil Shalin Panchal, and Daniel Kang · 2025
Closest in time.