Fetching the paper…
Reading the bibliography…
We are releasing a new suite of security benchmarks for LLMs, CYBERSECEVAL 3, to continue the conversation on empirically measuring LLM cybersecurity risks and capabilities.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Earlier work this paper cites.
Spear phishing with large language models, 2023
Julian Hazell · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Earlier work this paper cites.
Github copilot for business is now available, 2023
Microsoft · 2023
Earlier work this paper cites.
ChatGPT’s new code interpreter has giant security hole, allows hackers to steal your data, 2023
Avram Piltch · 2023
Earlier work this paper cites.
Voluntary AI commitments, 2023
United States White House · 2023
Earlier work this paper cites.
Multi-modal prompt injection image attacks against GPT-4V, 2023
Simon Willison · 2023
Earlier work this paper cites.
Mazal Bethany, Athanasios Galiopoulos, Emet Bethany, Mohammad Bahrami Karkevandi, Nishant Vishwamitra, and Peyman Najafirad · 2024
Earlier work this paper cites.
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024
Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, David Molnar, Spencer Whitman, and Joshua Saxe · 2024
Cited alongside, same era.
The near-term impact of AI on the cyber threat, 2024
UK National Cyber Security Centre · 2024
Cited alongside, same era.
eyeballvul: a future-proof benchmark for vulnerability detection in the wild, 2024
Timothee Chauvin · 2024
Cited alongside, same era.
Llm agents can autonomously exploit one-day vulnerabilities, 2024
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang · 2024
Cited alongside, same era.
Project naptime: Evaluating offensive security capabilities of large language models, 2024
Staying ahead of threat actors in the age of AI, 2024
Microsoft · 2024
Closest in time.
AI risk management framework, 2024
NIST · 2024
Closest in time.
OWASP Top Ten LLM Applications and Generative AI, 2024
OWASP · 2024
Closest in time.
No, LLM agents can not autonomously exploit one-day vulnerabilities, 2024
Chris Rohlf · 2024
Closest in time.
Runsybil web site, 2024
RunSybil · 2024
Closest in time.
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergei Glazunov and Mark Brand · 2024
Cited alongside, same era.
Generative AI for pentesting: the good, the bad, and the ugly
Eric Hilario, Sami Azar, Jawahar Sundaram, Kwaja Imran Mohamed, and Bharanidran Shanmugam · 2024
Cited alongside, same era.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models, 2024
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao · 2024
Cited alongside, same era.
Introducing XBOW, 2024
XBOW · 2024
Closest in time.