Fetching the paper…
Reading the bibliography…
Mainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2020 · 2004
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Qi, F.; Chen, Y.; Li, M.; Yao, Y.; Liu, Z.; and Sun, M. 2020 · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Biggio, B.; Nelson, B.; and Laskov, P. 2012 · 2012
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Gu, T.; Dolan-Gavitt, B.; and Garg, S. 2017 · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O.; and Savarese, S. 2017 · 2017
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2021 · 2021
Earlier work this paper cites.
Hidden killer: Invisible textual backdoor attacks with syntactic trigger
Qi, F.; Li, M.; Chen, Y.; Zhang, Z.; Liu, Z.; Wang, Y.; and Sun, M. 2021 · 2021
Earlier work this paper cites.
Chen, Y.; Gao, H.; Cui, G.; Qi, F.; Huang, L.; Liu, Z.; and Sun, M. 2022 · 2022
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue
Schulman, J.; Zoph, B.; Kim, C.; Hilton, J.; Menick, J.; Weng, J.; Uribe, J. F. C.; Fedus, L.; Metz, L.; Pokorny, M.; et al. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Poisoning web-scale training datasets is practical
Carlini, N.; Jagielski, M.; Choquette-Choo, C. A.; Paleka, D.; Pearce, W.; Anderson, H.; Terzis, A.; Thomas, K.; and Tramèr, F. 2023 · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023 · 2023
Cited alongside, same era.
Revolutionizing cyber threat detection with large language models
Ferrag, M. A.; Ndhlovu, M.; Tihanyi, N.; Cordeiro, L. C.; Debbah, M.; and Lestable, T. 2023 · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Universal jailbreak backdoors from poisoned human feedback
Rando, J.; and Tramèr, F. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
Cao, Y.; Cao, B.; and Chen, J. 2024 · 2024
Closest in time.
Stealthy Targeted Backdoor Attacks against Image Captioning
Fan, W.; Li, H.; Jiang, W.; Hao, M.; Yu, S.; and Zhang, X. 2024 · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Li, L.; Dong, B.; Wang, R.; Hu, X.; Zuo, W.; Lin, D.; Qiao, Y.; and Shao, J. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023 · 2023
Cited alongside, same era.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y.; Gupta, S.; Xia, M.; Li, K.; and Chen, D. 2023 · 2023
Cited alongside, same era.
SuRe: Improving Open-domain Question Answering of LLMs via Summarized Retrieval
Kim, J.; Nam, J.; Mo, S.; Park, J.; Lee, S.-W.; Seo, M.; Ha, J.-W.; and Shin, J. 2023 · 2023
Cited alongside, same era.
Efficient Ransomware Detection via Portable Executable File Image Analysis By LLaMA-7b
Li, X.; Zhu, T.; and Zhang, W. 2023 · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study
Liu, Y.; Deng, G.; Xu, Z.; Li, Y.; Zheng, Y.; Zhang, Y.; Zhao, L.; Zhang, T.; and Liu, Y. 2023 · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2023 · 2023
Cited alongside, same era.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Deng, G.; Liu, Y.; Li, Y.; Wang, K.; Zhang, Y.; Li, Z.; Wang, H.; Zhang, T.; and Liu, Y. 2023a
Cited in the paper.
Multilingual jailbreak challenges in large language models
Deng, Y.; Zhang, W.; Pan, S. J.; and Bing, L. 2023b
Cited in the paper.
Closest in time.
On the exploitability of instruction tuning
Shu, M.; Wang, J.; Zhu, C.; Geiping, J.; Xiao, C.; and Goldstein, T. 2024 · 2024
Closest in time.
Badchain: Backdoor chain-of-thought prompting for large language models
Xiang, Z.; Jiang, F.; Xiong, Z.; Ramasubramanian, B.; Poovendran, R.; and Li, B. 2024 · 2024
Closest in time.
Instruction backdoor attacks against customized { \{ LLMs } \}
Zhang, R.; Li, H.; Wen, R.; Jiang, W.; Zhang, Y.; Backes, M.; Shen, Y.; and Zhang, Y. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.
Toolqa: A dataset for llm question answering with external tools
Zhuang, Y.; Yu, Y.; Wang, K.; Sun, H.; and Zhang, C. 2024 · 2024
Closest in time.