Fetching the paper…
Reading the bibliography…
Backdoor attacks, in which a model behaves maliciously when given an attacker-specified trigger, pose a major security risk for practitioners who depend on publicly released language models.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The trojai software framework: An opensource tool for embedding trojans into deep learning models
Kiran Karra, Chace Ashcraft, and Neil Fendley. 2020 · 2003
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. 2017 · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David A. Wagner. 2018 · 2018
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona de Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros. 2018 · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018 · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018 · 2018
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019 · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao. 2019 · 2019
Earlier work this paper cites.
Universal litmus patterns: Revealing backdoor attacks in cnns
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann. 2020 · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
A tale of evil twins: Adversarial inputs versus poisoned models
Ren Pang, Hua Shen, Xinyang Zhang, Shouling Ji, Yevgeniy Vorobeychik, Xiapu Luo, Alex Liu, and Ting Wang. 2020 · 2020
Earlier work this paper cites.
T-Miner: A generative approach to defend against trojan attacks on DNN-based text classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K. Reddy, and Bimal Viswanath. 2021 · 2021
Earlier work this paper cites.
Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification
Chuanshuai Chen and Jiazhu Dai. 2021 · 2021
Earlier work this paper cites.
Trojan signatures in DNN weights
Greg Fields, Mohammad Samragh, Mojan Javaheripi, Farinaz Koushanfar, and Tara Javidi. 2021 · 2021
Cited alongside, same era.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Cited alongside, same era.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2021a · 2021
Cited alongside, same era.
Mind the style of text! adversarial and backdoor attacks based on text style transfer
Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, and Maosong Sun. 2021b · 2021
Cited alongside, same era.
Backdoor pre-trained models can transfer to all
Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang. 2021 · 2021
Mitigating backdoor poisoning attacks through the lens of spurious correlation
Xuanli He, Qiongkai Xu, Jun Wang, Benjamin Rubinstein, and Trevor Cohn. 2023 · 2023
Later among the works it cites.
Training-free lexical backdoor attacks on language models
Yujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu, Xingliang Yuan, and Chunyang Chen. 2023 · 2023
Later among the works it cites.
Tdc 2023 (llm edition): The trojan detection challenge
Mantas Mazeika, Andy Zou, Norman Mu, Long Phan, Zifan Wang, Chunru Yu, Adam Khoja, Fengqing Jiang, Aidan O’Gara, Ellie Sakhaee, Zhen Xiang, Arezoo Rajabi, Dan Hendrycks, Radha Poovendran, Bo Li, and David Forsyth. 2023c · 2023
Later among the works it cites.
Test-time backdoor mitigation for black-box large language models with defensive demonstrations
Wenjie Mo, Jiashu Xu, Qin Liu, Jiongxiao Wang, Jun Yan, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Later among the works it cites.
BITE: Textual backdoor attacks with iterative trigger injection
Jun Yan, Vansh Gupta, and Xiang Ren. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Detecting ai trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li. 2021 · 2021
Cited alongside, same era.
RAP: Robustness-Aware Perturbations for defending against backdoor attacks on NLP models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun. 2021a · 2021
Cited alongside, same era.
Expose backdoors on the way: A feature-based efficient defense against textual backdoor attacks
Sishuo Chen, Wenkai Yang, Zhiyuan Zhang, Xiaohan Bi, and Xu Sun. 2022b · 2022
Cited alongside, same era.
A unified evaluation of textual backdoor learning: Frameworks and benchmarks
Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Cited alongside, same era.
Planting undetectable backdoors in machine learning models : [extended abstract]
Shafi Goldwasser, Michael P. Kim, Vinod Vaikuntanathan, and Or Zamir. 2022 · 2022
Cited alongside, same era.
A study of the attention abnormality in trojaned BERTs
Weimin Lyu, Songzhu Zheng, Tengfei Ma, and Chao Chen. 2022 · 2022
Cited alongside, same era.
The trojan detection challenge
Mantas Mazeika, Dan Hendrycks, Huichen Li, Xiaojun Xu, Sidney Hough, Andy Zou, Arezoo Rajabi, Qi Yao, Zihao Wang, Jian Tian, Yao Tang, Di Tang, Roman Smirnov, Pavel Pleskov, Nikita Benkovich, Dawn Song, Radha Poovendran, Bo Li, and David. Forsyth. 2022 · 2022
Cited alongside, same era.
Gradient shaping: Enhancing backdoor attack against reverse engineering
Rui Zhu, Di Tang, Siyuan Tang, Guanhong Tao, Shiqing Ma, Xiaofeng Wang, and Haixu Tang. 2023 · 2023
Later among the works it cites.
Here‘s a free lunch: Sanitizing backdoored models with model merge
Ansh Arora, Xuanli He, Maximilian Mozes, Srinibas Swain, Mark Dras, and Qiongkai Xu. 2024 · 2024
Closest in time.
Alpagasus: Training a better alpaca with fewer data
Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al. 2024 · 2024
Closest in time.
From shortcuts to triggers: Backdoor defense with denoised PoE
Qin Liu, Fei Wang, Chaowei Xiao, and Muhao Chen. 2024 · 2024
Closest in time.
On model outsourcing adaptive attacks to deep learning backdoor defenses
Huaibing Peng, Huming Qiu, Hua Ma, Shuo Wang, Anmin Fu, Said F. Al-Sarawi, Derek Abbott, and Yansong Gao. 2024 · 2024
Closest in time.
On the (in)feasibility of ML backdoor detection as an hypothesis testing problem
Georg Pichler, Marco Romanelli, Divya Prakash Manivannan, Prashanth Krishnamurthy, Farshad khorrami, and Siddharth Garg. 2024 · 2024
Closest in time.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr. 2024 · 2024
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2024 · 2024
Closest in time.
Piccolo: Exposing complex backdoors in nlp transformer models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang. 2022 · 2042
Closest in time.