Fetching the paper…
Reading the bibliography…
With rapid advances, generative large language models (LLMs) dominate various Natural Language Processing (NLP) tasks from understanding to reasoning.
Weight poisoning attacks on pre-trained models
Kurita, K.; Michel, P.; and Neubig, G. 2020 · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; et al. 2020 · 2005
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Qi, F.; Chen, Y.; Li, M.; Yao, Y.; Liu, Z.; and Sun, M. 2020 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Towards Making Systems Forget with Machine Unlearning
Cao, Y.; and Yang, J. 2015 · 2015
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Chen, B.; Carvalho, W.; Baracaldo, N.; Ludwig, H.; Edwards, B.; Lee, T.; Molloy, I.; and Srivastava, B. 2018 · 2018
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Liu, K.; Dolan-Gavitt, B.; and Garg, S. 2018 · 2018
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Dai, J.; Chen, C.; and Li, Y. 2019 · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Gu, T.; Liu, K.; Dolan-Gavitt, B.; and Garg, S. 2019 · 2019
Earlier work this paper cites.
{ \{ T-Miner } \} : A generative approach to defend against trojan attacks on { \{ DNN-based } \} text classification
Azizi, A.; Tahmid, I. A.; Waheed, A.; Mangaokar, N.; Pu, J.; Javed, M.; Reddy, C. K.; and Viswanath, B. 2021 · 2021
Earlier work this paper cites.
Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification
Chen, C.; and Dai, J. 2021 · 2021
Earlier work this paper cites.
Badnl: Backdoor attacks against nlp models with semantic-preserving improvements
Chen, X.; Salem, A.; Chen, D.; Backes, M.; Ma, S.; Shen, Q.; Wu, Z.; and Zhang, Y. 2021 · 2021
Earlier work this paper cites.
Text backdoor detection using an interpretable rnn abstract model
Fan, M.; Si, Z.; Xie, X.; Liu, Y.; and Liu, T. 2021 · 2021
Earlier work this paper cites.
Triggerless backdoor attack for NLP tasks with clean labels
Gan, L.; Li, J.; Zhang, T.; Li, X.; Meng, Y.; Wu, F.; Yang, Y.; Guo, S.; and Fan, C. 2021 · 2021
Earlier work this paper cites.
Aligning AI With Shared Human Values
Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021a · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Li, Y.; Lyu, X.; Koren, N.; Lyu, L.; Li, B.; and Ma, X. 2021 · 2021
Earlier work this paper cites.
You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
Schuster, R.; Song, C.; Tromer, E.; and Shmatikov, V. 2021 · 2021
Earlier work this paper cites.
Concealed Data Poisoning Attacks on NLP Models
Wallace, E.; Zhao, T.; Feng, S.; and Singh, S. 2021 · 2021
Cited alongside, same era.
Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models
Yang, W.; Lin, Y.; Li, P.; Zhou, J.; and Sun, X. 2021 · 2021
Cited alongside, same era.
Spinning language models: Risks of propaganda-as-a-service and countermeasures
Bagdasaryan, E.; and Shmatikov, V. 2022 · 2022
Cited alongside, same era.
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Cited alongside, same era.
Kallima: A clean-label framework for textual backdoor attacks
Chen, X.; Dong, Y.; Sun, Z.; Zhai, S.; Shen, Q.; and Wu, Z. 2022 · 2022
Cited alongside, same era.
Backdoor Learning on Sequence to Sequence Models
Chen, L.; Cheng, M.; and Huang, H. 2023 · 2023
Later among the works it cites.
Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023 · 2023
Later among the works it cites.
Mitigating Backdoor Poisoning Attacks through the Lens of Spurious Correlation
He, X.; Xu, Q.; Wang, J.; Rubinstein, B.; and Cohn, T. 2023 · 2023
Later among the works it cites.
Composite backdoor attacks against large language models
Huang, H.; Zhao, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2023 · 2023
Later among the works it cites.
Knowledge Unlearning for Mitigating Privacy Risks in Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; et al. 2022 · 2022
Cited alongside, same era.
A Unified Evaluation of Textual Backdoor Learning: Frameworks and Benchmarks
Cui, G.; Yuan, L.; He, B.; Chen, Y.; Liu, Z.; and Sun, M. 2022 · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. 2022 · 2022
Cited alongside, same era.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Cited alongside, same era.
Backdoor Defense with Machine Unlearning
Liu, Y.; Fan, M.; Chen, C.; Liu, X.; Ma, Z.; Wang, L.; and Ma, J. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; et al. 2022 · 2022
Cited alongside, same era.
Hidden Trigger Backdoor Attack on NLP Models via Linguistic Style Manipulation
Pan, X.; Zhang, M.; Sheng, B.; Zhu, J.; and Yang, M. 2022 · 2022
Cited alongside, same era.
Jang, J.; Yoon, D.; Yang, S.; Cha, S.; Lee, M.; Logeswaran, L.; and Seo, M. 2023 · 2023
Later among the works it cites.
Multi-target Backdoor Attacks for Code Pre-trained Models
Li, Y.; Liu, S.; Chen, K.; Xie, X.; Zhang, T.; and Liu, Y. 2023 · 2023
Later among the works it cites.
OpenOrca: An Open Dataset of GPT Augmented FLAN Reasoning Traces
Lian, W.; Goodson, B.; Pentland, E.; Cook, A.; Vong, C.; and ”Teknium”. 2023 · 2023
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023 · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Rando, J.; and Tramèr, F. 2023 · 2023
Later among the works it cites.
Defending against backdoor attacks in natural language generation
Sun, X.; Li, X.; Meng, Y.; Ao, X.; Lyu, L.; Li, J.; and Zhang, T. 2023 · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; et al. 2023 · 2023
Later among the works it cites.
BITE: Textual Backdoor Attacks with Iterative Trigger Injection
Yan, J.; Gupta, V.; and Ren, X. 2023 · 2023
Later among the works it cites.
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Zhao, S.; Wen, J.; Tuan, L. A.; Zhao, J.; and Fu, J. 2023 · 2023
Later among the works it cites.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; et al. 2023 · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Later among the works it cites.
Sleeper agents: Training Deceptive LLMs that Persist Through Safety Training
Hubinger, E.; Denison, C.; Mu, J.; Lambert, M.; et al. 2024 · 2024
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Yan, J.; Yadav, V.; Li, S.; Chen, L.; Tang, Z.; Wang, H.; Srinivasan, V.; Ren, X.; and Jin, H. 2024 · 2024
Closest in time.
Latent backdoor attacks on deep neural networks
Yao, Y.; Li, H.; Zheng, H.; and Zhao, B. Y. 2019 · 2055
Closest in time.