Fetching the paper…
Reading the bibliography…
The safety defense methods of Large language models(LLMs) stays limited because the dangerous prompts are manually curated to just few known attack types, which fails to keep pace with emerging varieties.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2010
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Gpt-j-6b: A 6 billion parameter autoregressive language model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.
Diffusion-lm improves controllable text generation
Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. 2022 · 2022
Earlier work this paper cites.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Cited alongside, same era.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Cited alongside, same era.
Masterkey: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Cited alongside, same era.
Llm censorship: A machine learning challenge or a computer security problem?
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot, and Vardan Papyan. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral. 2023 · 2023
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang. 2023 · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Helbling, Mansi Phute, Matthew Hull, and Duen Horng Chau. 2023 · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Cited alongside, same era.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Soheil Feizi, and Hima Lakkaraju. 2023 · 2023
Cited alongside, same era.
Sok: Certified robustness for deep neural networks
Linyi Li, Tao Xie, and Bo Li. 2023 · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023 · 2023
Cited alongside, same era.
Mi Zhang, Xudong Pan, and Min Yang. 2023 · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Harnessing the plug-and-play controller by prompting
Hao Wang and Lei Sha. 2024 · 2024
Closest in time.
Gradient-based language model red teaming
Nevan Wichers, Carson Denison, and Ahmad Beirami. 2024 · 2024
Closest in time.