Fetching the paper…
Reading the bibliography…
Large language models (LLMs) introduce new security risks, but there are few comprehensive evaluation suites to measure and reduce these risks.
Bye bye bye: Evolution of repeated token attacks on chatgpt models, 2022
Mark Breitenbach and Adrian Wood · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Earlier work this paper cites.
Zefang Liu · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
https://owasp.org/Top10/A03_2021-Injection/ , 2021
Owasp top web application security threats · 2024
Cited alongside, same era.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Cited alongside, same era.
Considerations for evaluating large language models for cybersecurity tasks
Jeff Gennari, Shing-hon Lau, Samuel Perl, Joel Parish, and Girish Sastry · 2024
Cited alongside, same era.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Cited alongside, same era.
Cyberbench: A multi-task benchmark for evaluating large language models in cybersecurity
Zefang Liu, Jialei Shi, and John F Buford · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, et al · 2024
Closest in time.
Llm4vuln: A unified evaluation framework for decoupling and enhancing llms’ vulnerability reasoning
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Wei Ma, Lyuye Zhang, Miaolei Shi, and Yang Liu · 2024
Closest in time.
Cybermetric: A benchmark dataset for evaluating large language models knowledge in cybersecurity
Norbert Tihanyi, Mohamed Amine Ferrag, Ridhi Jain, and Merouane Debbah · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…