Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) like GPT-4, LLaMA, and Qwen have demonstrated remarkable success across a wide range of applications.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 1901
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti. 2020 · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Large-scale chemical language representations capture molecular structure and properties
Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das. 2022 · 2022
Earlier work this paper cites.
distilbert-prompt-injection
Deepset. 2023 · 2023
Earlier work this paper cites.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Earlier work this paper cites.
deberta-v3-base-injection
Fmops. 2023 · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023 · 2023
Earlier work this paper cites.
Hijacking context in large multi-modal models
Joonhyun Jeong. 2023 · 2023
Earlier work this paper cites.
Robust safety classifier for large language models: Adversarial prompt shield
Jinhwa Kim, Ali Derakhshan, and Ian G Harris. 2023 · 2023
Earlier work this paper cites.
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. 2023 · 2023
Earlier work this paper cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-AI conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. 2023 · 2023
Earlier work this paper cites.
Can chatgpt forecast stock price movements? return predictability and large language models
Alejandro Lopez-Lira and Yuehua Tang. 2023 · 2023
Earlier work this paper cites.
Designing chemical reaction arrays using phactor and chatgpt
Babak Mahjour, Jillian Hoffstadt, and Tim Cernak. 2023 · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Eloundou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Earlier work this paper cites.
Chatgpt: can artificial intelligence language models be of value for cardiovascular nurses and allied health professionals
Philip Moons and Liesbet Van Bulck. 2023 · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Cited alongside, same era.
Safety assessment of chinese large language models
Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023 · 2023
Cited alongside, same era.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024 · 2024
Closest in time.
Automatic and universal prompt injection attacks against large language models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024 · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. 2024 · 2024
Closest in time.
Prompt-guard-86m
meta.com. 2024 · 2024
Closest in time.
Atlas matrix
MITRE ATLAS. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kai-Cheng Yang and Filippo Menczer. 2023 · 2023
Cited alongside, same era.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Yiming Zhang and Daphne Ippolito. 2023 · 2023
Cited alongside, same era.
Defending large language models against jailbreaking attacks through goal prioritization
Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023 · 2023
Cited alongside, same era.
Chatgpt and environmental research
Jun-Jie Zhu, Jinyue Jiang, Meiqi Yang, and Zhiyong Jason Ren. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
alespalla/chatbot_instruction_prompts
Alespalla. 2023 · 2024
Cited alongside, same era.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al. 2024 · 2024
Cited alongside, same era.
Hyperion
Epivolis. 2024 · 2024
Cited alongside, same era.
Owasp top 10 list for large language models version 1.1
OWASP. 2024 · 2024
Closest in time.
Fine-tuned deberta-v3-base for prompt injection detection
ProtectAI.com. 2024 · 2024
Closest in time.
“Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024 · 2024
Closest in time.
Vmware/open-instruct
VMware. 2023 · 2024
Closest in time.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Closest in time.
Whylabs langkit
WhyLabs. 2024 · 2024
Closest in time.
Gradsafe: Detecting jailbreak prompts for llms via safety-critical gradient analysis
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong. 2024 · 2024
Closest in time.
A comprehensive study of jailbreak attack versus defense for large language models
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. 2024 · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al. 2024 · 2024
Closest in time.
Weak-to-strong jailbreaking on large language models
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang. 2024 · 2024
Closest in time.