Fetching the paper…
Reading the bibliography…
The rapid development of large language models (LLMs) has opened new avenues across various fields, including cybersecurity, which faces an evolving threat landscape and demand for innovative technologies.
Towards more accurate retrieval of duplicate bug reports
Chengnian Sun, David Lo, Siau-Cheng Khoo, and Jing Jiang · 2011
Earlier work this paper cites.
Standardizing cyber threat intelligence information with the structured threat information expression (stix)
Sean Barnum · 2012
Earlier work this paper cites.
An investigation on cyber security threats and security models
Kutub Thakur, Meikang Qiu, Keke Gai, and Md Liakat Ali · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
α \alpha diff: cross-version binary code similarity detection with dnn
Bingchang Liu, Wei Huo, Chao Zhang, Wenchao Li, Feng Li, Aihua Piao, and Wei Zou · 2018
Earlier work this paper cites.
A unified cybersecurity framework for complex environments
Jabu Mtsweni, Noluxolo Gcaza, and Mphahlele Thaba · 2018
Earlier work this paper cites.
Risk and the five hard problems of cybersecurity
Natalie M Scala, Allison C Reilly, Paul L Goethals, and Michel Cukier · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Earlier work this paper cites.
A comprehensive review study of cyber-attacks and cyber security; emerging trends and recent developments
Yuchong Li and Qinghui Liu · 2021
Earlier work this paper cites.
On the effectiveness of adapter-based tuning for pretrained language model adaptation
Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, and Luo Si · 2021
Earlier work this paper cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Ben Wang · 2021
Earlier work this paper cites.
Automatic program repair with openai’s codex: Evaluating quixbugs
Julian Aron Prenner and Romain Robbes · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Barret Zoph, Colin Raffel, Dale Schuurmans, Dani Yogatama, Denny Zhou, Don Metzler, Ed H Chi, Jason Wei, Jeff Dean, Liam B Fedus, et al · 2022
Earlier work this paper cites.
Cyber security, cyber threats, implications and future perspectives: A review
Diptiben Ghelani · 2022
Earlier work this paper cites.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna C. S. Santos · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
Jtrans: Jump-aware transformer for binary code similarity detection
Hao Wang, Wenjie Qu, Gilad Katz, Wenyu Zhu, Zeyu Gao, Han Qiu, Jianwei Zhuge, and Chao Zhang · 2022
Earlier work this paper cites.
Pop quiz! can a large language model help with reverse engineering?
Hammond Pearce, Benjamin Tan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, and Brendan Dolan-Gavitt · 2022
Earlier work this paper cites.
Practical program repair in the era of large pre-trained language models
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Earlier work this paper cites.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, et al · 2023
Earlier work this paper cites.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2023
Earlier work this paper cites.
Large language models in finance: A survey
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen · 2023
Earlier work this paper cites.
A comprehensive review of cyber security vulnerabilities, threats, attacks, and solutions
Ömer Aslan, Semih Serkant Aktuğ, Merve Ozkan-Okay, Abdullah Asim Yilmaz, and Erdal Akin · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Repairllama: Efficient representations and fine-tuned adapters for program repair
André Silva, Sen Fang, and Martin Monperrus · 2023
Earlier work this paper cites.
Hackmentor: Fine-tuning large language models for cybersecurity
Jie Zhang, Hui Wen, Liting Deng, Mingfeng Xin, Zhi Li, Lun Li, Hongsong Zhu, and Limin Sun · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, and et al · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Earlier work this paper cites.
Artificial intelligence for cybersecurity: Literature review and future research directions
Ramanpreet Kaur, Dušan Gabrijelčič, and Tomaž Klobučar · 2023
Earlier work this paper cites.
Artificial intelligence: revolutionizing cyber security in the digital era
Sarvesh Kumar, Upasana Gupta, Arvind Kumar Singh, and Avadh Kishore Singh · 2023
Earlier work this paper cites.
Towards artificial intelligence-based cybersecurity: the practices and chatgpt generated ways to combat cybercrime
Maad Mijwil, Mohammad Aljanabi, et al · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe · 2023
Earlier work this paper cites.
Llmseceval: A dataset of natural language prompts for security evaluations
Catherine Tony, Markus Mutas, Nicolás E. Díaz Ferreyra, and Riccardo Scandariato · 2023
Earlier work this paper cites.
Instruction tuning for large language models: A survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Earlier work this paper cites.
How abilities in large language models are affected by supervised fine-tuning data composition
Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou · 2023
Earlier work this paper cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Earlier work this paper cites.
Securefalcon: The next cyber reasoning system for cyber security
Mohamed Amine Ferrag, Ammar Battah, Norbert Tihanyi, Merouane Debbah, Thierry Lestable, and Lucas C Cordeiro · 2023
Earlier work this paper cites.
Seceval: A comprehensive benchmark for evaluating cybersecurity knowledge of foundation models
Guancheng Li, Yifeng Li, Wang Guannan, Haoyu Yang, and Yang Yu · 2023
Earlier work this paper cites.
Secqa: A concise question-answering dataset for evaluating large language models in computer security
Zefang Liu · 2023
Earlier work this paper cites.
An empirical study of netops capability of pre-trained large language models
Yukai Miao, Yu Bai, Li Chen, Dan Li, Haifeng Sun, Xizheng Wang, Ziqiu Luo, Yanyu Ren, Dapeng Sun, Xiuting Xu, Qi Zhang, Chao Xiang, and Xinchi Li · 2023
Earlier work this paper cites.
Baichuan 2: Open large-scale language models
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, et al · 2023
Earlier work this paper cites.
Gpt understands, too
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2023
Earlier work this paper cites.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang · 2023
Earlier work this paper cites.
Generative ai and prompt engineering: The art of whispering to let the genie out of the algorithmic world
Aras Bozkurt and Ramesh C Sharma · 2023
Earlier work this paper cites.
Prompt engineering a prompt engineer
Qinyuan Ye, Maxamed Axmed, Reid Pryzant, and Fereshte Khani · 2023
Earlier work this paper cites.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2023
Earlier work this paper cites.
Codegen2: Lessons for training llms on programming and natural languages
Erik Nijkamp, Hiroaki Hayashi, Caiming Xiong, Silvio Savarese, and Yingbo Zhou · 2023
Earlier work this paper cites.
Efficient avoidance of vulnerabilities in auto-completed smart contract code using vulnerability-constrained decoding
André Storhaug, Jingyue Li, and Tianyuan Hu · 2023
Earlier work this paper cites.
Nova + : Generative language models for binaries
Nan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu, Lin Tan, and Xiangyu Zhang · 2023
Earlier work this paper cites.
AGIR: automating cyber threat intelligence reporting with natural language generation
Filippo Perrina, Francesco Marchiori, Mauro Conti, and Nino Vincenzo Verde · 2023
Earlier work this paper cites.
On the uses of large language models to interpret ambiguous cyberattack descriptions
Reza Fayyazi and Shanchieh Jay Yang · 2023
Earlier work this paper cites.
An empirical study on using large language models to analyze software supply chain security failures
Tanmay Singla, Dharun Anandayuvaraj, Kelechi G. Kalu, Taylor R. Schorlemmer, and James C. Davis · 2023
Earlier work this paper cites.
Time for action: Automated analysis of cyber threat intelligence in the wild
Giuseppe Siracusano, Davide Sanvito, Roberto Gonzalez, Manikantan Srinivasan, Sivakaman Kamatchi, Wataru Takahashi, Masaru Kawakita, Takahiro Kakumaru, and Roberto Bifulco · 2023
Earlier work this paper cites.
Cupid: Leveraging chatgpt for more accurate duplicate bug report detection
Ting Zhang, Ivana Clairine Irsan, Ferdian Thung, and David Lo · 2023
Earlier work this paper cites.
Hw-v2w-map: Hardware vulnerability to weakness mapping framework for root cause analysis with gpt-assisted mitigation suggestion
Yu-Zheng Lin, Muntasir Mamun, Muhtasim Alam Chowdhury, Shuyu Cai, Mingyu Zhu, Banafsheh Saber Latibari, Kevin Immanuel Gubbi, Najmeh Nazari Bavarsad, Arjun Caputo, Avesta Sasan, Houman Homayoun, Setareh Rafatirad, Pratik Satam, and Soheil Salehi · 2023
Earlier work this paper cites.
Cyber sentinel: Exploring conversational agents in streamlining security tasks with gpt-4
Mehrdad Kaheh, Danial Khosh Kholgh, and Panos Kostakos · 2023
Earlier work this paper cites.
Evaluation of chatgpt model for vulnerability detection
Anton Cheshkov, Pavel Zadorozhny, and Rodion Levichev · 2023
Earlier work this paper cites.
Software vulnerability detection using large language models
Moumita Das Purba, Arpita Ghosh, Benjamin J. Radford, and Bill Chu · 2023
Earlier work this paper cites.
Vuldetect: A novel technique for detecting software vulnerabilities using language models
Marwan Omar and Stavros Shiaeles · 2023
Earlier work this paper cites.
Understanding the effectiveness of large language models in detecting security vulnerabilities
Avishree Khare, Saikat Dutta, Ziyang Li, Alaia Solko-Breslin, Rajeev Alur, and Mayur Naik · 2023
Earlier work this paper cites.
The hitchhiker’s guide to program analysis: A journey with large language models
Haonan Li, Yu Hao, Yizhuo Zhai, and Zhiyun Qian · 2023
Earlier work this paper cites.
Defecthunter: A novel llm-driven boosted-conformer-based code vulnerability detection mechanism
Jin Wang, Zishan Huang, Hengli Liu, Nianyi Yang, and Yinhao Xiao · 2023
Earlier work this paper cites.
Using chatgpt as a static application security testing tool
Atieh Bakhshandeh, Abdalsamad Keramatfar, Amir Norouzi, and Mohammad Mahdi Chekidehkhoun · 2023
Earlier work this paper cites.
Large language model-powered smart contract vulnerability detection: New perspectives
Sihao Hu, Tiansheng Huang, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2023
Earlier work this paper cites.
Software vulnerability detection with gpt and in-context learning
Zhihong Liu, Qing Liao, Wenchao Gu, and Cuiyun Gao · 2023
Earlier work this paper cites.
Vullibgen: Identifying vulnerable third-party libraries via generative pre-trained model
Tianyu Chen, Lin Li, Liuchuan Zhu, Zongyang Li, Guangtai Liang, Ding Li, Qianxiang Wang, and Tao Xie · 2023
Earlier work this paper cites.
How chatgpt is solving vulnerability management problem
Peiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu, Yifan Xia, Xuhong Zhang, Wenzhi Chen, Haiqin Weng, Shouling Ji, and Wenhai Wang · 2023
Earlier work this paper cites.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner · 2023
Earlier work this paper cites.
How far have we gone in vulnerability detection using large language models
Zeyu Gao, Hao Wang, Yuchen Zhou, Wenyu Zhu, and Chao Zhang · 2023
Earlier work this paper cites.
The formai dataset: Generative ai in software security through the lens of formal verification
Norbert Tihanyi, Tamas Bisztray, Ridhi Jain, Mohamed Amine Ferrag, Lucas C. Cordeiro, and Vasileios Mavroeidis · 2023
Earlier work this paper cites.
Understanding programs by exploiting (fuzzing) test cases
Jianyu Zhao, Yuyang Rong, Yiwen Guo, Yifeng He, and Hao Chen · 2023
Earlier work this paper cites.
Evaluating and explaining large language models for code using syntactic structures
David N Palacio, Alejandro Velasco, Daniel Rodriguez-Cardenas, Kevin Moran, and Denys Poshyvanyk · 2023
Earlier work this paper cites.
Prompt engineering-assisted malware dynamic analysis using gpt-4
Pei Yan, Shunquan Tan, Miaohui Wang, and Jiwu Huang · 2023
Earlier work this paper cites.
Using chatgpt to analyze ransomware messages and to predict ransomware threats, 2023
Himari Fujima, Takako Kumamoto, and Yunko Yoshida · 2023
Earlier work this paper cites.
Using large language models to mitigate ransomware threats
Fang Wang · 2023
Earlier work this paper cites.
Flag: Finding line anomalies (in code) with generative ai
Baleegh Ahmad, Benjamin Tan, Ramesh Karri, and Hammond Pearce · 2023
Earlier work this paper cites.
Log-based anomaly detection based on evt theory with feedback
Jinyang Liu, Junjie Huang, Yintong Huo, Zhihan Jiang, Jiazhen Gu, Zhuangbin Chen, Cong Feng, Minzhi Yan, and Michael R. Lyu · 2023
Earlier work this paper cites.
Loggpt: Exploring chatgpt for log-based anomaly detection
Jiaxing Qi, Shaohan Huang, Zhongzhi Luan, Shu Yang, Carol J. Fung, Hailong Yang, Depei Qian, Jing Shang, Zhiwen Xiao, and Zhihui Wu · 2023
Earlier work this paper cites.
Loggpt: Log anomaly detection via GPT
Xiao Han, Shuhan Yuan, and Mohamed Trabelsi · 2023
Earlier work this paper cites.
An improved transformer-based model for detecting phishing, spam, and ham: A large language model approach
Suhaima Jamal and Hayden Wimmer · 2023
Earlier work this paper cites.
Devising and detecting phishing: Large language models vs. smaller human models
Fredrik Heiding, Bruce Schneier, Arun Vishwanath, Jeremy Bernstein, and Peter S. Park · 2023
Earlier work this paper cites.
Web content filtering through knowledge distillation of large language models
Tamás Vörös, Sean Paul Bergeron, and Konstantin Berlin · 2023
Earlier work this paper cites.
Application of large language models to ddos attack detection
Michael Guastalla, Yiyi Li, Arvin Hekmati, and Bhaskar Krishnamachari · 2023
Earlier work this paper cites.
Explaining tree model decisions in natural language for network intrusion detection
Noah Ziems, Gang Liu, John Flanagan, and Meng Jiang · 2023
Earlier work this paper cites.
Huntgpt: Integrating machine learning-based anomaly detection and explainable ai with large language models (llms)
Tarek Ali and Panos Kostakos · 2023
Earlier work this paper cites.
Chatgpt for digital forensic investigation: The good, the bad, and the unknown
Mark Scanlon, Frank Breitinger, Christopher Hargreaves, Jan-Niclas Hilgert, and John Sheppard · 2023
Earlier work this paper cites.
How well does llm generate security tests?
Ying Zhang, Wenjia Song, Zhengjie Ji, Danfeng, Yao, and Na Meng · 2023
Earlier work this paper cites.
Augmenting greybox fuzzing with generative ai
Jie Hu, Qian Zhang, and Heng Yin · 2023
Earlier work this paper cites.
Codamosa: Escaping coverage plateaus in test generation ·with pre-trained large language models
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, and Siddhartha Sen · 2023
Earlier work this paper cites.
Understanding large language model based fuzz driver generation
Cen Zhang, Mingqiang Bai, Yaowen Zheng, Yeting Li, Xiaofei Xie, Yuekang Li, Wei Ma, Limin Sun, and Yang Liu · 2023
Earlier work this paper cites.
Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models
Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang · 2023
Earlier work this paper cites.
Large language models are edge-case fuzzers: Testing deep learning libraries via fuzzgpt
Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang · 2023
Earlier work this paper cites.
An analysis of the automatic bug fixing performance of chatgpt
Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke · 2023
Earlier work this paper cites.
Examining zero-shot vulnerability repair with large language models
Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt · 2023
Earlier work this paper cites.
How effective are neural networks for fixing security vulnerabilities
Yi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan, Petr Babkin, and Sameena Shah · 2023
Earlier work this paper cites.
Inferfix: End-to-end program repair with llms
Matthew Jin, Syed Shahriar, Michele Tufano, Xin Shi, Shuai Lu, Neel Sundaresan, and Alexey Svyatkovskiy · 2023
Earlier work this paper cites.
Better patching using llm prompting, via self-consistency
Toufique Ahmed and Premkumar Devanbu · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Earlier work this paper cites.
Copiloting the copilots: Fusing large language models with completion engines for automated program repair
Yuxiang Wei, Chunqiu Steven Xia, and Lingming Zhang · 2023
Earlier work this paper cites.
Zeroleak: Using llms for scalable and cost effective side-channel patching
M. Caner Tol and Berk Sunar · 2023
Earlier work this paper cites.
Divas: An llm-based end-to-end framework for soc security analysis and policy-based protection
Sudipta Paria, Aritra Dasgupta, and Swarup Bhunia · 2023
Earlier work this paper cites.
Fixing hardware security bugs with large language models
Baleegh Ahmad, Shailja Thakur, Benjamin Tan, Ramesh Karri, and Hammond Pearce · 2023
Earlier work this paper cites.
Identifying and mitigating the security risks of generative ai
Clark Barrett, Brad Boyd, Elie Bursztein, Nicholas Carlini, Brad Chen, Jihye Choi, Amrita Roy Chowdhury, Mihai Christodorescu, Anupam Datta, Soheil Feizi, et al · 2023
Cited alongside, same era.
Impact of big data analytics and chatgpt on cybersecurity
Pawankumar Sharma and Bibhu Dash · 2023
Cited alongside, same era.
From chatgpt to threatgpt: Impact of generative ai in cybersecurity and privacy
Maanak Gupta, Charankumar Akiri, Kshitiz Aryal, Eli Parker, and Lopamudra Praharaj · 2023
Cited alongside, same era.
Llms killed the script kiddie: How agents supported by large language models change the landscape of network threat testing
Stephen Moskal, Sam Laney, Erik Hemberg, and Una-May O’Reilly · 2023
Cited alongside, same era.
Pentestgpt: An llm-empowered automatic penetration testing tool
Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass · 2023
Cited alongside, same era.
Harnessing the power of llms in source code vulnerability detection
Andrew A Mahyari · 2024
Closest in time.
Towards effectively detecting and explaining vulnerabilities using large language models
Qiheng Mao, Zhenhao Li, Xing Hu, Kui Liu, Xin Xia, and Jianling Sun · 2024
Closest in time.
Software vulnerability and functionality assessment using llms
Rasmus Ingemann Tuffveson Jensen, Vali Tawosi, and Salwa Alamir · 2024
Closest in time.
Assessing the effectiveness of llms in android application vulnerability analysis
Vasileios Kouliaridis, Georgios Karopoulos, and Georgios Kambourakis · 2024
Closest in time.
Outside the comfort zone: Analysing llm capabilities in software vulnerability detection
Yuejun Guo, Constantinos Patsakis, Qiang Hu, Qiang Tang, and Fran Casino · 2024
Closest in time.
Prompt-enhanced software vulnerability detection using chatgpt
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Getting pwn’d by ai: Penetration testing with large language models
Andreas Happe and Jürgen Cito · 2023
Cited alongside, same era.
Penheal: A two-stage llm framework for automated pentesting and optimal remediation
Junjie Huang and Quanyan Zhu · 2023
Cited alongside, same era.
Exploring the dark side of ai: Advanced phishing attack design and deployment using chatgpt
Nils Begou, Jérémy Vinoy, Andrzej Duda, and Maciej Korczyński · 2023
Cited alongside, same era.
Evaluating llms for privilege-escalation scenarios
Andreas Happe, Aaron Kaplan, and Jürgen Cito · 2023
Cited alongside, same era.
From text to mitre techniques: Exploring the malicious use of large language models for generating cyber attack payloads
P. V. Sai Charan, Hrushikesh Chunduri, P. Mohan Anand, and Sandeep K Shukla · 2023
Cited alongside, same era.
Using large language models for cybersecurity capture-the-flag challenges and certification questions
Wesley Tann, Yuancheng Liu, Jun Heng Sim, Choon Meng Seah, and Ee-Chien Chang · 2023
Cited alongside, same era.
Ratgpt: Turning online llms into proxies for malware attacks
Mika Beckerich, Laura Plein, and Sergio Coronado · 2023
Cited alongside, same era.
Chenyuan Zhang, Hao Liu, Jiutian Zeng, Kejing Yang, Yuhong Li, and Hui Li · 2024
Closest in time.
Llbezpeky: Leveraging large language models for vulnerability detection
Noble Saji Mathews, Yelizaveta Brus, Yousra Aafer, Mei Nagappan, and Shane McIntosh · 2024
Closest in time.
Static detection of filesystem vulnerabilities in android systems
Yu-Tsung Lee, Hayawardh Vijayakumar, Zhiyun Qian, and Trent Jaeger · 2024
Closest in time.
Anvil: Anomaly-based vulnerability identification without labelled training data
Weizhou Wang, Eric Liu, Xiangyu Guo, and David Lie · 2024
Closest in time.
Simclone: Detecting tabular data clones using value similarity
Xu Yang, Gopi Krishnan Rajbahadur, Dayi Lin, Shaowei Wang, and Zhen Ming (Jack) Jiang · 2024
Closest in time.
Vul-rag: Enhancing llm-based vulnerability detection via knowledge-level rag
Xueying Du, Geng Zheng, Kaixin Wang, Jiayi Feng, Wentai Deng, Mingwei Liu, Bihuan Chen, Xin Peng, Tao Ma, and Yiling Lou · 2024
Closest in time.
Gptscan: Detecting logic vulnerabilities in smart contracts by combining GPT with program analysis
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Haijun Wang, Zhengzi Xu, Xiaofei Xie, and Yang Liu · 2024
Closest in time.
Generalization-enhanced code vulnerability detection via multi-task instruction fine-tuning
Xiaohu Du, Ming Wen, Jiahao Zhu, Zifan Xie, Bin Ji, Huijun Liu, Xuanhua Shi, and Hai Jin · 2024
Closest in time.
Llm4vuln: A unified evaluation framework for decoupling and enhancing llms’ vulnerability reasoning
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Wei Ma, Lyuye Zhang, Miaolei Shi, and Yang Liu · 2024
Closest in time.
Multi-role consensus through llms discussions for vulnerability detection
Zhenyu Mao, Jialong Li, Munan Li, and Kenji Tei · 2024
Closest in time.
Llm-assisted static analysis for detecting security vulnerabilities
Ziyang Li, Saikat Dutta, and Mayur Naik · 2024
Closest in time.
Scope: Evaluating llms for software vulnerability detection
José Gonçalves, Tiago Dias, Eva Maia, and Isabel Praça · 2024
Closest in time.
Llm4decompile: Decompiling binary code with large language models
Hanzhuo Tan, Qi Luo, Jing Li, and Yuqun Zhang · 2024
Closest in time.
Large language models for code analysis: Do llms really do their job?
Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu, Ruoyu Zhang, Ruijie Fang, Asmita, Ryan Tsang, Najmeh Nazari, Han Wang, and Houman Homayoun · 2024
Closest in time.
Shifting the lens: Detecting malware in npm ecosystem with large language models
Nusrat Zahan, Philipp Burckhardt, Mikola Lysenko, Feross Aboukhadijeh, and Laurie Williams · 2024
Closest in time.
Malsight: Exploring malicious source code and benign pseudocode for iterative binary malware summarization
Haolang Lu, Hongrui Peng, Guoshun Nan, Jiaoyang Cui, Cheng Wang, Weifei Jin, Songtao Wang, Shengli Pan, and Xiaofeng Tao · 2024
Closest in time.
Make LLM a testing expert: Bringing human-like interaction to mobile GUI testing via functionality-aware decisions
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang · 2024
Closest in time.
Benchmarking large language models for log analysis, security, and interpretation
Egil Karlsen, Xiao Luo, Nur Zincir-Heywood, and Malcolm I. Heywood · 2024
Closest in time.
Interpretable online log analysis using large language models with prompt strategies
Yilun Liu, Shimin Tao, Weibin Meng, Jingyu Wang, Wenbing Ma, Yuhang Chen, Yanqing Zhao, Hao Yang, and Yanfei Jiang · 2024
Closest in time.
Lemur: Log parsing with entropy sampling and chain-of-thought merging
Wei Zhang, Hongcheng Guo, Anjie Le, Jian Yang, Jiaheng Liu, Zhoujun Li, Tieqiao Zheng, Shi Xu, Runqiang Zang, Liangfan Zheng, and Bo Zhang · 2024
Closest in time.
Evaluating the performance of chatgpt for spam email detection
Yuwei Wu, Shijing Si, Yugui Zhang, Jiawen Gu, and Jedrek Wosik · 2024
Closest in time.
Prompted contextual vectors for spear-phishing detection
Daniel Nahmias, Gal Engelberg, Dan Klein, and Asaf Shabtai · 2024
Closest in time.
Revolutionizing cyber threat detection with large language models: A privacy-preserving bert-based lightweight model for iot/iiot devices
Mohamed Amine Ferrag, Mthandazo Ndhlovu, Norbert Tihanyi, Lucas C. Cordeiro, Mérouane Debbah, Thierry Lestable, and Narinderjit Singh Thandi · 2024
Closest in time.
When fuzzing meets llms: Challenges and opportunities
Yu Jiang, Jie Liang, Fuchen Ma, Yuanliang Chen, Chijin Zhou, Yuheng Shen, Zhiyong Wu, Jingzhou Fu, Mingzhe Wang, Shanshan Li, et al · 2024
Closest in time.
An exploratory study on using large language models for mutation testing
Bo Wang, Mingda Chen, Youfang Lin, Mike Papadakis, and Jie M Zhang · 2024
Closest in time.
Fuzz4all: Universal fuzzing with large language models
Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang · 2024
Closest in time.
Large language model guided protocol fuzzing
Ruijie Meng, Martin Mirchev, Marcel Böhme, and Abhik Roychoudhury · 2024
Closest in time.
Fuzzing busybox: Leveraging LLM and crash reuse for embedded bug unearthing
Asmita, Yaroslav Oliinyk, Michael Scott, Ryan Tsang, Chongzhou Fang, and Houman Homayoun · 2024
Closest in time.
A systematic literature review on large language models for automated program repair
Quanjun Zhang, Chunrong Fang, Yang Xie, YuXiang Ma, Weisong Sun, Yun Yang, and Zhenyu Chen · 2024
Closest in time.
Ai-powered patching: the future of automated vulnerability fixes
Jan Keller and Jan Nowakowski · 2024
Closest in time.
Security code review by llms: A deep dive into responses
Jiaxin Yu, Peng Liang, Yujia Fu, Amjed Tahir, Mojtaba Shahin, Chong Wang, and Yangxiao Cai · 2024
Closest in time.
How far can we go with practical function-level program repair?
Jiahong Xiang, Xiaoyang Xu, Fanchu Kong, Mingyuan Wu, Zizheng Zhang, Haotian Zhang, and Yuqun Zhang · 2024
Closest in time.
Aligning llms for fl-free program repair
Junjielong Xu, Ying Fu, Shin Hwei Tan, and Pinjia He · 2024
Closest in time.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou · 2024
Closest in time.
A case study of llm for automated vulnerability repair: Assessing impact of reasoning and patch validation feedback
Ummay Kulsum, Haotian Zhu, Bowen Xu, and Marcelo d’Amorim · 2024
Closest in time.
Enhancing automated program repair with solution design
Jiuang Zhao, Donghao Yang, Li Zhang, Xiaoli Lian, Zitian Yang, and Fang Liu · 2024
Closest in time.
Revisiting unnaturalness for automated program repair in the era of large language models
Aidan Z. H. Yang, Sophia Kolak, Vincent J. Hellendoorn, Ruben Martins, and Claire Le Goues · 2024
Closest in time.
Thinkrepair: Self-directed automated program repair
Xin Yin, Chao Ni, Shaohua Wang, Zhenhao Li, Limin Zeng, and Xiaohu Yang · 2024
Closest in time.
Llm-powered code vulnerability repair with reinforcement learning and semantic reward
Nafis Tanveer Islam, Joseph Khoury, Andrew Seong, Mohammad Bahrami Karkevandi, Gonzalo De La Torre Parra, Elias Bou-Harb, and Peyman Najafirad · 2024
Closest in time.
Revisiting evolutionary program repair via code language model
Yunan Wang, Tingyu Guo, Zilong Huang, and Yuan Yuan · 2024
Closest in time.
Contrastrepair: Enhancing conversation-based automated program repair via contrastive test case pairs
Jiaolong Kong, Mingfei Cheng, Xiaofei Xie, Shangqing Liu, Xiaoning Du, and Qi Guo · 2024
Closest in time.
When large language models confront repository-level automatic program repair: How well they done?
Yuxiao Chen, Jingzheng Wu, Xiang Ling, Changjiang Li, Zhiqing Rui, Tianyue Luo, and Yanjun Wu · 2024
Closest in time.
Multi-objective fine-tuning for enhanced program repair with llms
Boyang Yang, Haoye Tian, Jiadong Ren, Hongyu Zhang, Jacques Klein, Tegawendé F. Bissyandé, Claire Le Goues, and Shunfu Jin · 2024
Closest in time.
Enhanced automated code vulnerability repair using large language models
David de-Fitero-Dominguez, Eva García-López, Antonio García-Cabot, and José Javier Martínez-Herráiz · 2024
Closest in time.
Repair: Automated program repair with process-based feedback
Yuze Zhao, Zhenya Huang, Yixiao Ma, Rui Li, Kai Zhang, Hao Jiang, Qi Liu, Linbo Zhu, and Yu Su · 2024
Closest in time.
Mergerepair: An exploratory study on merging task-specific adapters in code llms for automated program repair
Meghdad Dehghan, Jie JW Wu, Fatemeh H. Fard, and Ali Ouni · 2024
Closest in time.
Automated repair of ai code with large language models and formal verification
Yiannis Charalambous, Edoardo Manino, and Lucas C. Cordeiro · 2024
Closest in time.
A study of vulnerability repair in javascript programs with large language models
Tan Khang Le, Saba Alimadadi, and Steven Y. Ko · 2024
Closest in time.
Automated c/c++ program repair for high-level synthesis via large language models
Kangwei Xu, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, and Bing Li · 2024
Closest in time.
Malla: Demystifying real-world large language model integrated malicious services
Zilong Lin, Jian Cui, Xiaojing Liao, and XiaoFeng Wang · 2024
Closest in time.
Cipher: Cybersecurity intelligent penetration-testing helper for ethical researcher
Derry Pratama, Naufal Suryanto, Andro Aprila Adiputra, Thi-Thu-Huong Le, Ahmada Yusril Kadiptya, Muhammad Iqbal, and Howon Kim · 2024
Closest in time.
Autoattacker: A large language model guided system to implement automatic cyber-attacks
Jiacen Xu, Jack W. Stokes, Geoff McDonald, Xuesong Bai, David Marshall, Siyue Wang, Adith Swaminathan, and Zhou Li · 2024
Closest in time.
From sands to mansions: Enabling automatic full-life-cycle cyberattack construction with llm
Lingzhi Wang, Jiahui Wang, Kyle Jung, Kedar Thiagarajan, Emily Wei, Xiangmin Shen, Yan Chen, and Zhenyuan Li · 2024
Closest in time.
Is generative ai the next tactical cyber weapon for threat actors? unforeseen implications of ai generated cyber attacks
Yusuf Usman, Aadesh Upadhyay, Prashnna Gyawali, and Robin Chataut · 2024
Closest in time.
From chatbots to phishbots? – preventing phishing scams created using chatgpt, google bard and claude
Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, and Shirin Nilizadeh · 2024
Closest in time.
Assessing ai vs human-authored spear phishing sms attacks: An empirical study using the trapd method
Jerson Francia, Derek Hansen, Ben Schooley, Matthew Taylor, Shydra Murray, and Greg Snow · 2024
Closest in time.
Using retriever augmented large language models for attack graph generation
Renascence Tarafder Prapty, Ashish Kundu, and Arun Iyengar · 2024
Closest in time.
Bugs in large language models generated code: An empirical study
Florian Tambon, Arghavan Moradi Dakhel, Amin Nikanjam, Foutse Khomh, Michel C. Desmarais, and Giuliano Antoniol · 2024
Closest in time.
Do neutral prompts produce insecure code? formai-v2 dataset: Labelling vulnerabilities in code generated by large language models
Norbert Tihanyi, Tamas Bisztray, Mohamed Amine Ferrag, Ridhi Jain, and Lucas C. Cordeiro · 2024
Closest in time.
Is your ai-generated code really safe? evaluating large language models on secure code generation with codeseceval
Jiexin Wang, Xitong Luo, Liuwen Cao, Hongkui He, Hailin Huang, Jiayuan Xie, Adam Jatowt, and Yi Cai · 2024
Closest in time.
No need to lift a finger anymore? assessing the quality of code generation by chatgpt
Zhijie Liu, Yutian Tang, Xiapu Luo, Yuming Zhou, and Liang Feng Zhang · 2024
Closest in time.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2024
Closest in time.
Llm security guard for code
Arya Kavian, Mohammad Mehdi Pourhashem Kallehbasti, Sajjad Kazemi, Ehsan Firouzi, and Mohammad Ghafari · 2024
Closest in time.
An exploratory study on fine-tuning large language models for secure code generation
Junjie Li, Fazle Rabbi, Cheng Cheng, Aseem Sangalay, Yuan Tian, and Jinqiu Yang · 2024
Closest in time.
Code repair with llms gives an exploration-exploitation tradeoff
Hao Tang, Keya Hu, Jin Peng Zhou, Sicheng Zhong, Wei-Long Zheng, Xujie Si, and Kevin Ellis · 2024
Closest in time.
Investigating the transferability of code repair for low-resource programming languages
Kyle Wong, Alfonso Amayuelas, Liangming Pan, and William Yang Wang · 2024
Closest in time.
LLM for soc security: A paradigm shift
Dipayan Saha, Shams Tarek, Katayoon Yahyaei, Sujan Kumar Saha, Jingbo Zhou, Mark M. Tehranipoor, and Farimah Farahmandi · 2024
Closest in time.
LLM in the shell: Generative honeypots
Muris Sladic, Veronica Valeros, Carlos Adrián Catania, and Sebastian García · 2024
Closest in time.
Act as a honeytoken generator! an investigation into honeytoken generation with large language models
Daniel Reti, Norman Becker, Tillmann Angeli, Anasuya Chattopadhyay, Daniel Schneider, Sebastian Vollmer, and Hans D. Schotten · 2024
Closest in time.
Llmpot: Automated llm-based industrial protocol and physical process emulation for ics honeypots
Christoforos Vasilatos, Dunia J. Mahboobeh, Hithem Lamri, Manaar Alam, and Michail Maniatakos · 2024
Closest in time.
Employing llms for incident response planning and review
Sam Hays and Dr. Jules White · 2024
Closest in time.
Prompting is all you need: Automated android bug replay with large language models
Sidong Feng and Chunyang Chen · 2024
Closest in time.
Is stack overflow obsolete? an empirical study of the characteristics of chatgpt answers to stack overflow questions
Samia Kabir, David N. Udo-Imeh, Bonan Kou, and Tianyi Zhang · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang · 2024
Closest in time.
Strengthening llm trust boundaries: A survey of prompt injection attacks surender suresh kumar dr. ml cummings dr. alexander stimpson
Surender Suresh Kumar, ML Cummings, and Alexander Stimpson · 2024
Closest in time.
A new era in llm security: Exploring security concerns in real-world llm-based systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, and Chaowei Xiao · 2024
Closest in time.
Comprehensive assessment of jailbreak attacks against llms
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang · 2024
Closest in time.
Llm jailbreak attack versus defense techniques–a comprehensive study
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek · 2024
Closest in time.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
Shuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan, and Jinming Wen · 2024
Closest in time.
Poisonprompt: Backdoor attack on prompt-based large language models
Hongwei Yao, Jian Lou, and Zhan Qin · 2024
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2024
Closest in time.
Jatmo: Prompt injection defense by task-specific finetuning
Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David A. Wagner · 2024
Closest in time.
A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2024
Closest in time.
Masterkey: Automated jailbreaking of large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2024
Closest in time.
Autodan: interpretable gradient-based adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
POSTER: identifying and mitigating vulnerabilities in llm-integrated applications
Fengqing Jiang, Zhangchen Xu, Luyao Niu, Boxin Wang, Jinyuan Jia, Bo Li, and Radha Poovendran · 2024
Closest in time.
Breaking the silence: the threats of using llms in software engineering
June Sallou, Thomas Durieux, and Annibale Panichella · 2024
Closest in time.
Llmind: Orchestrating ai and iot with llm for complex task execution
Hongwei Cui, Yuyang Du, Qun Yang, Yulin Shao, and Soung Chang Liew · 2024
Closest in time.
Out of the cage: How stochastic parrots win in cyber security environments
Maria Rigaki, Ondrej Lukás, Carlos Adrián Catania, and Sebastian García · 2024
Closest in time.
Llm agents can autonomously hack websites
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Closest in time.
Nissist: An incident mitigation copilot based on troubleshooting guides
Kaikai An, Fangkai Yang, Junting Lu, Liqun Li, Zhixing Ren, Hao Huang, Lu Wang, Pu Zhao, Yu Kang, Hua Ding, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Qi Zhang · 2024
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
From summary to action: Enhancing large language models for complex tasks with open world apis
Yulong Liu, Yunlong Yuan, Chunwei Wang, Jianhua Han, Yongqiang Ma, Li Zhang, Nanning Zheng, and Hang Xu · 2024
Closest in time.
If llm is the wizard, then code is the wand: A survey on how code empowers large language models to serve as intelligent agents
Ke Yang, Jiateng Liu, John Wu, Chaoqi Yang, Yi R. Fung, Sha Li, Zixuan Huang, Xu Cao, Xingyao Wang, Yiquan Wang, Heng Ji, and Chengxiang Zhai · 2024
Closest in time.
Llm agents can autonomously exploit one-day vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang · 2024
Closest in time.
Teams of llm agents can exploit zero-day vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Closest in time.
Phishagent: A robust multimodal agent for phishing webpage detection
Tri Cao, Chengyu Huang, Yuexin Li, Huilin Wang, Amy He, Nay Oo, and Bryan Hooi · 2024
Closest in time.
Using llms to automate threat intelligence analysis workflows in security operation centers
PeiYu Tseng, ZihDwo Yeh, Xushu Dai, and Peng Liu · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for LLM agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Closest in time.
Wipi: A new web threat for llm-driven web agents
Fangzhou Wu, Shutong Wu, Yulong Cao, and Chaowei Xiao · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang · 2024
Closest in time.
Evaluation of llm-based chatbots for osint-based cyber threat awareness
Samaneh Shafee, Alysson Bessani, and Pedro M. Ferreira · 2025
Closest in time.
Dlap: A deep learning augmented large language model prompting framework for software vulnerability detection
Yanjing Yang, Xin Zhou, Runfeng Mao, Jinwei Xu, Lanxin Yang, Yu Zhang, Haifeng Shen, and He Zhang · 2025
Closest in time.