Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) are deployed and integrated into thousands of applications, the need for scalable evaluation of how models respond to adversarial attacks grows rapidly.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Limits of static analysis for malware detection
Andreas Moser, Christopher Kruegel, and Engin Kirda. 2007 · 2007
Earlier work this paper cites.
Fuzzing: brute force vulnerability discovery
Michael Sutton, Adam Greene, and Pedram Amini. 2007 · 2007
Earlier work this paper cites.
Jigsaw unintended bias in toxicity classification
cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Security Engineering: A guide to building dependable distributed systems
Ross Anderson. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Earlier work this paper cites.
Vulnerability disclosure mechanisms: A synthesis and framework for market-based and non-market-based disclosures
Ali Ahmed, Amit Deokar, and Ho Cheung Brian Lee. 2021 · 2021
Earlier work this paper cites.
An idr framework of opportunities and barriers between hci and nlp
Nanna Inie and Leon Derczynski. 2021 · 2021
Earlier work this paper cites.
Red Teaming Handbook (3rd Edition)
UK Ministry of Defence. 2021 · 2021
Earlier work this paper cites.
The surprising performance of simple baselines for misinformation detection
Kellin Pelrine, Jacob Danovitch, and Reihaneh Rabbany. 2021 · 2021
Earlier work this paper cites.
Ai and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021 · 2021
Earlier work this paper cites.
Process for adapting language models to society (palms) with values-targeted datasets
Irene Solaiman and Christy Dennison. 2021 · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022 · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
The internal state of an LLM knows when it’s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Earlier work this paper cites.
Katie Moussouris: Vulnerability Disclosure and Security Workforce Development
Bob Blakley and Lorrie Cranor. 2023 · 2023
Earlier work this paper cites.
Speak, memory: An archaeology of books known to ChatGPT/GPT-4
Kent K Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Cited alongside, same era.
Assessing Language Model Deployment with Risk Cards
Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yulia Tsvetkov, MR Leiser, and Saif Mohammad. 2023 · 2023
Cited alongside, same era.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Kaustubh Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahadiran, Simon Mille, Ashish Shrivastava, Samson Tan, et al. 2023 · 2023
Cited alongside, same era.
Closed AI Models Make Bad Baselines
Anna Rogers. 2023 · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Later among the works it cites.
OWASP Top 10 for LLM
Steve Wilson, Ads Dawson, Leon Derczynski, Mike Finch, Itamar Golan, Kai Greshake, Rich Harang, Ken Huang, Gavin Klondike, Autumn Moulder, Eugene Neelou, David Rowe, Manjesh S, Andy Smith, Rachit Sood, John Sotiropoulos, Andrew Amaro, Stefano Amorelli, Ken Arora, Jason Axley, Aliaksei Bialko, Patrick Biyaga, Larry Carson, Adrian Culley, Lior Drihem, Andy Dyrcz, Guillaume Ehinger, Vladimir Fedotov, Dan Frommer, Adesh Gairola, Cassio Goldschmidt, Nipun Gupta, Jason Haddix, Nathan Hamiel, Idan Hen, Bajram Hoxha, Mike Jang, Emmanuel Guilherme Junior, Dan Klein, Ananda Krishna, Santosh Kumar, Kelvin Low, Vishwas Manral, Matteo Große-Kampmann, Brodie McRae, Ross Moore, Dotan Nahum, Joshua Nussbaum, Gaurav “GP” Pal, Priyadharshini Parthasarathy, Nir Paz, Brian Pendleton, Jorge Pinto, James Rabe, Ashish Rajan, Reza Rashidi, Johann Rehberger, Jason Ross, Aleksei Ryzhkov, Talesh Seeparsan, Vandana Verma Sehgal, and Leonardo Shikida. 2023 · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. 2023 · 2023
Cited alongside, same era.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence
Executive Order 14110. 2023 · 2023
Cited alongside, same era.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. 2023 · 2023
Cited alongside, same era.
How We Broke LLMs: Indirect Prompt Injection
Kai Greshake. 2023 · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Cited alongside, same era.
Large language models can be used to effectively scale spear phishing campaigns
Julian Hazell. 2023 · 2023
Cited alongside, same era.
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the Wild
Nanna Inie, Jonathan Stray, and Leon Derczynski. 2023 · 2023
Cited alongside, same era.
Can you trust ChatGPT’s package recommendations?
Bar Lanyado. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Bing chat: Data exfiltration exploit explained
wunderwuzzi. 2023 · 2023
Later among the works it cites.
Security challenges in natural language processing models
Qiongkai Xu and Xuanli He. 2023 · 2023
Later among the works it cites.
GPTfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Nmap Introduction - Phrack 51, Article 11 — nmap.org
Fyodor. 1997 · 2024
Closest in time.
ISC2 Report: The Number of Women in Cybersecurity Remains Stagnant, Despite Ongoing Workforce Gap — asisonline.org
Megan Gates. 2024 · 2024
Closest in time.
PoC: LLM prompt injection via invisible instructions in pasted text
Riley Goodside. 2024 · 2024
Closest in time.
Glitch tokens in large language models: Categorization taxonomy and effective detection
Yuxi Li, Yi Liu, Gelei Deng, Ying Zhang, Wenjia Song, Ling Shi, Kailong Wang, Yuekang Li, Yang Liu, and Haoyu Wang. 2024 · 2024
Closest in time.
Against The Achilles’ Heel: A Survey on Red Teaming for Generative Models
Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang, Renxi Wang, Junjie Gao, Yixuan Zhang, Wanxiang Che, Timothy Baldwin, et al. 2024 · 2024
Closest in time.
Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
Apostol Vassilev, Alina Oprea, Alie Fordyce, and Hyrum Anderson. 2024 · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Closest in time.
Do-not-answer: Evaluating safeguards in LLMs
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2024a · 2024
Closest in time.