Adversarial stylometry in the wild: Transferable lexical substitution attacks on author profiling
Chris Emmery, Ákos Kádár, and Grzegorz Chrupała · 2021
Later among the works it cites.
Defending against backdoor attacks in natural language generation
Chun Fan, Xiaoya Li, Yuxian Meng, Xiaofei Sun, Xiang Ao, Fei Wu, Jiwei Li, and Tianwei Zhang · 2021
Later among the works it cites.
Model extraction and adversarial transferability, your BERT is vulnerable!
Xuanli He, Lingjuan Lyu, Qiongkai Xu, and Lichao Sun · 2021
Later among the works it cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu · 2021
Later among the works it cites.
Using adversarial attacks to reveal the statistical bias in machine reading comprehension models
Jieyu Lin, Jiajie Zou, and Nai Ding · 2021
Later among the works it cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Later among the works it cites.
Generating natural language attacks in a hard label black box setting
Rishabh Maheshwary, Saket Maheshwary, and Vikram Pudi · 2021
Later among the works it cites.
Adv-OLM: Generating textual adversaries via OLM
Vijit Malik, Ashwani Bhat, and Ashutosh Modi · 2021
Later among the works it cites.
Dataset reconstruction attack against language models
Rrubaa Panchendrarajan and Suman Bhoi · 2021
Later among the works it cites.
ONION: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Later among the works it cites.
Backdoor pre-trained models can transfer to all
Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang · 2021
Later among the works it cites.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
Colin Wei, Sang Michael Xie, and Tengyu Ma · 2021
Later among the works it cites.
Pre-trained models: Past, present and future
Han Xu, Zhang Zhengyan, Ding Ning, Gu Yuxian, Liu Xiao, Huo Yuqi, Qiu Jiezhong, Zhang Liang, Han Wentao, Huang Minlie, et al · 2021
Later among the works it cites.
Beyond model extraction: Imitation attack for black-box NLP APIs
Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, and Gholamreza Haffari · 2021
Later among the works it cites.
On the transferability of adversarial attacks against neural text classifier
Liping Yuan, Xiaoqing Zheng, Yi Zhou, Cho-Jui Hsieh, and Kai-Wei Chang · 2021
Later among the works it cites.
Grey-box extraction of natural language models
Santiago Zanella-Beguelin, Shruti Tople, Andrew Paverd, and Boris Köpf · 2021
Later among the works it cites.
Trojaning language models for fun and profit
Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang · 2021
Later among the works it cites.
Red alarm for pre-trained models: Universal vulnerability to neuron-level backdoor attacks
Zhengyan Zhang, Guangxuan Xiao, Yongwei Li, Tian Lv, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Xin Jiang, and Maosong Sun · 2021
Later among the works it cites.