Fetching the paper…
Reading the bibliography…
The landscape of adversarial attacks against text classifiers continues to grow, with new attacks developed every year and many of them available in standard toolkits, such as TextAttack and OpenAttack.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Natural language adversarial attacks and defenses in word level
Xiaosen Wang, Hao Jin, and Kun He. 2019 · 1909
Earlier work this paper cites.
Learning robust representations of text
Yitong Li, Trevor Cohn, and Timothy Baldwin. 2016 · 1985
Earlier work this paper cites.
Robustness verification for transformers
Zhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang, and Cho-Jui Hsieh. 2020 · 2002
Earlier work this paper cites.
Trends in spam products and methods
Geoff Hulten, Anthony Penta, Gopalakrishnan Seshadrinathan, and Manav Mishra. 2004 · 2004
Earlier work this paper cites.
Can machine learning be secure?
Marco Barreno, Blaine Nelson, Russell Sears, Anthony D. Joseph, and J. D. Tygar. 2006 · 2006
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
The enemy in your own camp: How well can we detect statistically-generated fake reviews – an adversarial study
Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
LightGBM: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Earlier work this paper cites.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M. Dai, and Ian Goodfellow. 2017 · 2017
Earlier work this paper cites.
Ex Machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018 · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018 · 2018
Earlier work this paper cites.
Source generator attribution via inversion
Michael Albright and Scott McCloskey. 2019 · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Text processing like humans do: Visually attacking and shielding NLP systems
Steffen Eger, Gözde Gül Şahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Achieving verified robustness to symbol substitutions via interval bound propagation
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli. 2019 · 2019
Cited alongside, same era.
A robust adversarial training approach to machine reading comprehension
Kai Liu, Xin Liu, An Yang, Jing Liu, Jinsong Su, Sujian Li, and Qiaoqiao She. 2020b · 2020
Later among the works it cites.
TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Later among the works it cites.
Mind your inflections! Improving NLP for non-standard English with base-inflection encoding
Samson Tan, Shafiq Joty, Lav R Varshney, and Min-Yen Kan. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
SAFER: A structure-free approach for certified robustness to adversarial word substitutions
Mao Ye, Chengyue Gong, and Qiang Liu. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 2019
Cited alongside, same era.
TextBugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019 · 2019
Cited alongside, same era.
Robust to noise models in natural language processing tasks
Valentin Malykh. 2019 · 2019
Cited alongside, same era.
Combating adversarial misspellings with robust word recognition
Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Word-level textual adversarial attacking as combinatorial optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. 2020 · 2020
Later among the works it cites.
FreeLB: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Towards robustness against natural language word substitutions
Xinshuai Dong, Anh Tuan Luu, Rongrong Ji, and Hong Liu. 2021 · 2021
Later among the works it cites.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021 · 2021
Later among the works it cites.
Achieving model robustness through discrete adversarial training
Maor Ivgi and Jonathan Berant. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
A sweet rabbit hole by DARCY: Using honeypots to detect universal trigger’s adversarial attacks
Thai Le, Noseong Park, and Dongwon Lee. 2021 · 2021
Later among the works it cites.
Searching for an effective defender: Benchmarking defense against adversarial word substitution
Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang, Kai-Wei Chang, and Cho-Jui Hsieh. 2021b · 2021
Later among the works it cites.
Sample efficient detection and classification of adversarial attacks via self-supervised embeddings
Mazda Moayeri and Soheil Feizi. 2021 · 2021
Later among the works it cites.
Frequency-guided word substitutions for detecting textual adversarial examples
Maximilian Mozes, Pontus Stenetorp, Bennett Kleinberg, and Lewis Griffin. 2021 · 2021
Later among the works it cites.
What models know about their attackers: Deriving attacker information from latent representations
Zhouhang Xie, Jonathan Brophy, Adam Noack, Wencong You, Kalyani Asthana, Carter Perkins, Sabrina Reis, Zayd Hammoudeh, Daniel Lowd, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Towards improving adversarial training of NLP models
Jin Yong Yoo and Yanjun Qi. 2021 · 2021
Later among the works it cites.
OpenAttack: An open-source textual adversarial attack toolkit
Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Later among the works it cites.