Fetching the paper…
Reading the bibliography…
Adversarial attack serves as a major challenge for neural network models in NLP, which precludes the model's deployment in safety-critical applications.
Towards a robust deep neural network in texts: A survey
Wenqi Wang, Run Wang, Lina Wang, Zhibo Wang, and Aoshuang Ye · 1902
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Nltk: the natural language toolkit
Edward Loper and Steven Bird · 2002
Earlier work this paper cites.
Yiming Li, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2007
Earlier work this paper cites.
On the negativity of negation
Christopher Potts · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
J Springenberg, Alexey Dosovitskiy, Thomas Brox, and M Riedmiller · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
Neural trojans
Yuntao Liu, Yang Xie, and Ankur Srivastava · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk · 2018
Earlier work this paper cites.
Adversarial examples for natural language classification problems
Volodymyr Kuleshov, Shantanu Thakoor, Tingfung Lau, and Stefano Ermon · 2018
Cited alongside, same era.
Semantically equivalent adversarial rules for debugging nlp models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Text processing like humans do: Visually attacking and shielding nlp systems
Steffen Eger, Gözde Gül Şahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych · 2019
Cited alongside, same era.
Achieving verified robustness to symbol substitutions via interval bound propagation
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli · 2019
Cited alongside, same era.
Word-level textual adversarial attacking as combinatorial optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun · 2020
Later among the works it cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li · 2020
Later among the works it cites.
Defending pre-trained language models from adversarial word substitution without performance sacrifice
Rongzhou Bao, Jiayi Wang, and Hai Zhao · 2021
Later among the works it cites.
Libre: A practical bayesian approach to adversarial detection
Zhijie Deng, Xiao Yang, Shizhen Xu, Hang Su, and Jun Zhu · 2021
Later among the works it cites.
Towards robustness against natural language word substitutions
Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu · 2021
Later among the works it cites.
Model extraction and adversarial transferability, your bert is vulnerable!
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang · 2019
Cited alongside, same era.
Popqorn: Quantifying robustness of recurrent neural networks
Ching-Yun Ko, Zhaoyang Lyu, Lily Weng, Luca Daniel, Ngai Wong, and Dahua Lin · 2019
Cited alongside, same era.
Detection based defense against adversarial examples from the steganalysis point of view
Jiayang Liu, Weiming Zhang, Yiwei Zhang, Dongdong Hou, Yujia Liu, Hongyue Zha, and Nenghai Yu · 2019
Cited alongside, same era.
Detecting and diagnosing adversarial images with class-conditional capsule reconstructions
Yao Qin, Nicholas Frosst, Sara Sabour, Colin Raffel, Garrison Cottrell, and Geoffrey Hinton · 2019
Cited alongside, same era.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che · 2019
Cited alongside, same era.
Robustness verification for transformers
Zhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Cited alongside, same era.
Xuanli He, Lingjuan Lyu, Lichao Sun, and Qiongkai Xu · 2021
Later among the works it cites.
A sweet rabbit hole by darcy: Using honeypots to detect universal trigger’s adversarial attacks
Thai Le, Noseong Park, and Dongwon Lee · 2021
Later among the works it cites.
Contextualized perturbation for textual adversarial attack
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan · 2021
Later among the works it cites.
Using adversarial attacks to reveal the statistical bias in machine reading comprehension models
Jieyu Lin, Jiajie Zou, and Nai Ding · 2021
Later among the works it cites.
Frequency-guided word substitutions for detecting textual adversarial examples
Maximilian Mozes, Pontus Stenetorp, Bennett Kleinberg, and Lewis Griffin · 2021
Later among the works it cites.
Crafting adversarial examples for neural machine translation
Xinze Zhang, Junzhe Zhang, Zhenhua Chen, and Kun He · 2021
Later among the works it cites.
Defense against adversarial attacks in nlp via dirichlet neighborhood ensemble
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-wei Chang, and Xuanjing Huang · 2021
Later among the works it cites.
A survey in adversarial defences and robustness in nlp
Shreya Goyal, Sumanth Doddapaneni, Mitesh M Khapra, and Balaraman Ravindran · 2022
Later among the works it cites.
Perturbations in the wild: Leveraging human-written text perturbations for realistic adversarial attack and defense
Thai Le, Jooyoung Lee, Kevin Yen, Yifan Hu, and Dongwon Lee · 2022
Later among the works it cites.
Flooding-x: Improving bert’s resistance to adversarial attacks via loss-restricted fine-tuning
Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, Zhihua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, and Xuan-Jing Huang · 2022
Later among the works it cites.
“that is a suspicious reaction!”: Interpreting logits variation to detect nlp adversarial attacks
Edoardo Mosca, Shreyash Agarwal, Javier Rando Ramírez, and Georg Groh · 2022
Later among the works it cites.
Semattack: Natural textual attacks via different semantic spaces
Boxin Wang, Chejian Xu, Xiangyu Liu, Yu Cheng, and Bo Li · 2022
Later among the works it cites.