Fetching the paper…
Reading the bibliography…
In the field of natural language processing, the prevalent approach involves fine-tuning pretrained language models (PLMs) using local samples.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Trojaning attack on neural networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2018
Earlier work this paper cites.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhilu Zhang and Mert Sabuncu · 2018
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 2019
Earlier work this paper cites.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar · 2019
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
An embarrassingly simple approach for trojan attack in deep neural networks
Ruixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang, and Xia Hu · 2020
Earlier work this paper cites.
Interpretability and analysis in neural nlp
Yonatan Belinkov, Sebastian Gehrmann, and Ellie Pavlick · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer · 2020
Earlier work this paper cites.
Mind the style of text! adversarial and backdoor attacks based on text style transfer
Fanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li, Zhiyuan Liu, and Maosong Sun · 2021
Earlier work this paper cites.
Hidden killer: Invisible textual backdoor attacks with syntactic trigger
Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun · 2021
Cited alongside, same era.
Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification
Chuanshuai Chen and Jiazhu Dai · 2021
Cited alongside, same era.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Cited alongside, same era.
Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun · 2021
Cited alongside, same era.
Design and evaluation of a multi-domain trojan detection method on deep neural networks
Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith C Ranasinghe, and Hyoungshick Kim · 2021
Cited alongside, same era.
Few-shot backdoor attacks on visual object tracking
Yiming Li, Haoxiang Zhong, Xingjun Ma, Yong Jiang, and Shu-Tao Xia · 2022
Later among the works it cites.
Piccolo: Exposing complex backdoors in nlp transformer models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Later among the works it cites.
Constrained optimization with dynamic bound-scaling for effective nlp backdoor defense
Guangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu, Zhuo Zhang, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2022
Later among the works it cites.
A study of the attention abnormality in trojaned berts
Weimin Lyu, Songzhu Zheng, Tengfei Ma, and Chao Chen · 2022
Later among the works it cites.
Moderate-fitting as a natural backdoor defender for pre-trained language models
Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Chong Fu, Yangdong Deng, Zhiyuan Liu, Jingang Wang, and Wei Wu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Detecting ai trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li · 2021
Cited alongside, same era.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2021
Cited alongside, same era.
Fairness via representation neutralization
Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang, Ahmed Awadallah, and Xia Hu · 2021
Cited alongside, same era.
Generating syntactically controlled paraphrases without using annotated parallel pairs
Kuan-Hao Huang and Kai-Wei Chang · 2021
Cited alongside, same era.
Anti-backdoor learning: Training clean models on poisoned data
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma · 2021
Cited alongside, same era.
Measure and improve robustness in nlp models: A survey
Xuezhi Wang, Haohan Wang, and Diyi Yang · 2022
Cited alongside, same era.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2022
Cited alongside, same era.
Trap and replace: Defending backdoor attacks by trapping them into an easy-to-replace subnetwork
Haotao Wang, Junyuan Hong, Aston Zhang, Jiayu Zhou, and Zhangyang Wang · 2022
Later among the works it cites.
Expose backdoors on the way: A feature-based efficient defense against textual backdoor attacks
Sishuo Chen, Wenkai Yang, Zhiyuan Zhang, Xiaohan Bi, and Xu Sun · 2022
Later among the works it cites.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2023
Closest in time.
Not all samples are born equal: Towards effective clean-label backdoor attacks
Yinghua Gao, Yiming Li, Linghui Zhu, Dongxian Wu, Yong Jiang, and Shu-Tao Xia · 2023
Closest in time.
Backdoor learning for nlp: Recent advances, challenges, and future research directions
Marwan Omar · 2023
Closest in time.
Chatgpt as an attack tool: Stealthy textual backdoor attack via blackbox generative model trigger
Jiazhao Li, Yijin Yang, Zhuofeng Wu, VG Vydiswaran, and Chaowei Xiao · 2023
Closest in time.
Untargeted backdoor attack against object detection
Chengxiao Luo, Yiming Li, Yong Jiang, and Shu-Tao Xia · 2023
Closest in time.
Label poisoning is all you need
Rishi Jha, Jonathan Hayase, and Sewoong Oh · 2023
Closest in time.
Large language models can be lazy learners: Analyze shortcuts in in-context learning
Ruixiang Tang, Dehan Kong, Longtao Huang, and Hui Xue · 2023
Closest in time.
Revisiting the assumption of latent separability for backdoor defenses
Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, and Prateek Mittal · 2023
Closest in time.