Fetching the paper…
Reading the bibliography…
Adversarial attacks are a major challenge faced by current machine learning research.
Detecting adversarial examples and other misclassifications in neural networks by introspection
Jonathan Aigrain and Marcin Detyniecki. 2019 · 1905
Earlier work this paper cites.
Natural language adversarial attacks and defenses in word level
Xiaosen Wang, Hao Jin, and Kun He. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020 · 1910
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Support vector machines
M.A. Hearst, S.T. Dumais, E. Osuna, J. Platt, and B. Scholkopf. 1998 · 1998
Earlier work this paper cites.
A brief introduction to boosting
Robert E. Schapire. 1999 · 1999
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee. 2005 · 2005
Earlier work this paper cites.
Defense against adversarial attacks in nlp via dirichlet neighborhood ensemble
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-wei Chang, and Xuanjing Huang. 2020 · 2006
Earlier work this paper cites.
Detection defense against adversarial attacks with saliency map
Dengpan Ye, Chuanxi Chen, Changrui Liu, Hao Wang, and Shunzhi Jiang. 2020 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Early methods for detecting adversarial images
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Earlier work this paper cites.
Adversarial examples in the physical world
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2017 · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Synthesizing robust adversarial examples
Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. 2018 · 2018
Cited alongside, same era.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019 · 2019
Later among the works it cites.
The odds are odd: A statistical test for detecting adversarial examples
Kevin Roth, Yannic Kilcher, and Thomas Hofmann. 2019 · 2019
Later among the works it cites.
A study on single and multi-layer perceptron neural network
Jaswinder Singh and Rajdeep Banerjee. 2019 · 2019
Later among the works it cites.
Learning to discriminate perturbations for blocking adversarial attacks in text classification
Yichao Zhou, Jyun-Yu Jiang, Kai-Wei Chang, and Wei Wang. 2019 · 2019
Later among the works it cites.
When explainability meets adversarial learning: Detecting adversarial examples using shap signatures
G. Fidel, R. Bitton, and A. Shabtai. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Black-box generation of adversarial text sequences to evade deep learning classifiers
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018 · 2018
Cited alongside, same era.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Harini Kannan, Alexey Kurakin, and Ian Goodfellow. 2018 · 2018
Cited alongside, same era.
Towards robust detection of adversarial examples
Tianyu Pang, Chao Du, Yinpeng Dong, and Jun Zhu. 2018 · 2018
Cited alongside, same era.
Attacks meet interpretability: Attribute-steered detection of adversarial samples
Guanhong Tao, Shiqing Ma, Yingqi Liu, and Xiangyu Zhang. 2018 · 2018
Cited alongside, same era.
Towards mitigating adversarial texts
Basemah Alshemali and Jugal Kalita. 2019 · 2019
Cited alongside, same era.
Siddhant Garg and Goutham Ramakrishnan. 2020 · 2020
Later among the works it cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Later among the works it cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020 · 2020
Later among the works it cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Later among the works it cites.
On adaptive attacks to adversarial example defenses
Florian Tramer, Nicholas Carlini, Wieland Brendel, and Aleksander Madry. 2020 · 2020
Later among the works it cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z. Sheng, Ahoud Alhazmi, and Chenliang Li. 2020 · 2020
Later among the works it cites.
Towards robustness against natural language word substitutions
Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu. 2021 · 2021
Later among the works it cites.
Understanding and interpreting the impact of user context in hate speech detection
Edoardo Mosca, Maximilian Wich, and Georg Groh. 2021 · 2021
Later among the works it cites.
Frequency-guided word substitutions for detecting textual adversarial examples
Maximilian Mozes, Pontus Stenetorp, Bennett Kleinberg, and Lewis Griffin. 2021 · 2021
Later among the works it cites.
Model-agnostic adversarial example detection through logit distribution learning
Yaopeng Wang, Lehui Xie, Ximeng Liu, Jia-Li Yin, and Tingjie Zheng. 2021 · 2021
Later among the works it cites.
Explainable abusive language classification leveraging user and network data
Maximilian Wich, Edoardo Mosca, Adrian Gorniak, Johannes Hingerl, and Georg Groh. 2021 · 2021
Later among the works it cites.