Fetching the paper…
Reading the bibliography…
The language models, especially the basic text classification models, have been shown to be susceptible to textual adversarial attacks such as synonym substitution and word insertion attacks.
J. Neyman and E. S. Pearson, “Ix. on the problem of the most efficient tests of statistical hypotheses,” Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character , vol. 231, no. 694-706, pp. 289–337, 1933
1933
Earlier work this paper cites.
“Bidirectional recurrent neural networks,” IEEE transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies , 2011, pp. 142–150
2011
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
S. Shalev-Shwartz et al. , “Online learning and online convex optimization,” Foundations and Trends® in Machine Learning , vol. 4, no. 2, pp. 107–194, 2012
2012
Earlier work this paper cites.
J. McAuley and J. Leskovec, “Hidden factors and hidden topics: understanding rating dimensions with review text,” in Proceedings of the 7th ACM conference on Recommender systems , 2013, pp. 165–172
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
Q. Geng, P. Kairouz, S. Oh, and P. Viswanath, “The staircase mechanism in differential privacy,” IEEE Journal of Selected Topics in Signal Processing , vol. 9, no. 7, pp. 1176–1184, 2015
2015
Earlier work this paper cites.
Y. Chen, “Convolutional neural network for sentence classification,” Master’s thesis, University of Waterloo, 2015
2015
Earlier work this paper cites.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Feng, E. Wallace, A. Grissom II, P. Rodriguez, M. Iyyer, and J. Boyd-Graber, “Pathologies of neural models make interpretation difficult,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , 2018, pp. 3719–3728
2018
Earlier work this paper cites.
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou, “Hotflip: White-box adversarial examples for text classification,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2018, pp. 31–36
2018
Earlier work this paper cites.
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi, “Black-box generation of adversarial text sequences to evade deep learning classifiers,” in 2018 IEEE Security and Privacy Workshops (SPW) . IEEE, 2018, pp. 50–56
2018
Earlier work this paper cites.
M. Iyyer, J. Wieting, K. Gimpel, and L. Zettlemoyer, “Adversarial example generation with syntactically controlled paraphrase networks,” in Proceedings of NAACL-HLT , 2018, pp. 1875–1885
2018
Earlier work this paper cites.
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. St. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, B. Strope, and R. Kurzweil, “Universal sentence encoder for English,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , E. Blanco and W. Lu, Eds. Association for Computational Linguistics, 2018, pp. 169–174
2018
Earlier work this paper cites.
M. Alzantot, Y. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang, “Generating natural language adversarial examples,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2018
2018
Earlier work this paper cites.
S. Ren, Y. Deng, K. He, and W. Che, “Generating natural language adversarial examples through probability weighted word saliency,” in Proceedings of the 57th annual meeting of the association for computational linguistics , 2019, pp. 1085–1097
2019
Earlier work this paper cites.
L. Wu, F. Morstatter, K. M. Carley, and H. Liu, “Misinformation in social media: definition, manipulation, and detection,” ACM SIGKDD explorations newsletter , vol. 21, no. 2, pp. 80–90, 2019
2019
Earlier work this paper cites.
J. Cohen, E. Rosenfeld, and Z. Kolter, “Certified adversarial robustness via randomized smoothing,” in international conference on machine learning . PMLR, 2019, pp. 1310–1320
2019
Earlier work this paper cites.
R. Jia, A. Raghunathan, K. Göksel, and P. Liang, “Certified robustness to adversarial word substitutions,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 4129–4142
2019
Earlier work this paper cites.
C.-Y. Ko, Z. Lyu, L. Weng, L. Daniel, N. Wong, and D. Lin, “Popqorn: Quantifying robustness of recurrent neural networks,” in International Conference on Machine Learning . PMLR, 2019, pp. 3468–3477
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
J. Li, S. Ji, T. Du, B. Li, and T. Wang, “Textbugger: Generating adversarial text against real-world applications,” in 26th Annual Network and Distributed System Security Symposium , 2019
2019
Cited alongside, same era.
D. Pruthi, B. Dhingra, and Z. C. Lipton, “Combating adversarial misspellings with robust word recognition,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5582–5591
2019
Cited alongside, same era.
M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, and S. Jana, “Certified robustness to adversarial examples with differential privacy,” in SP . IEEE, 2019, pp. 656–672
2019
Cited alongside, same era.
Y. Nie, Y. Wang, and M. Bansal, “Analyzing compositionality-sensitivity of nli models,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 6867–6874
2019
Cited alongside, same era.
X. Wang, J. Hao, Y. Yang, and K. He, “Natural language adversarial defense through synonym encoding,” in Uncertainty in Artificial Intelligence . PMLR, 2021, pp. 823–833
2021
Later among the works it cites.
T. Du, S. Ji, L. Shen, Y. Zhang, J. Li, J. Shi, C. Fang, J. Yin, R. Beyah, and T. Wang, “Cert-rnn: Towards certifying the robustness of recurrent neural networks.” CCS , vol. 21, no. 2021, pp. 15–19, 2021
2021
Later among the works it cites.
G. Bonaert, D. I. Dimitrov, M. Baader, and M. Vechev, “Fast and precise certification of transformers,” in Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation , 2021, pp. 466–481
2021
Later among the works it cites.
W. Wang, P. Tang, J. Lou, and L. Xiong, “Certified robustness to word substitution attack with differential privacy,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 1102–1112
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Behjati, S.-M. Moosavi-Dezfooli, M. S. Baghshah, and P. Frossard, “Universal adversarial attacks on text classifiers,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7345–7349
2019
Cited alongside, same era.
P.-S. Huang, R. Stanforth, J. Welbl, C. Dyer, D. Yogatama, S. Gowal, K. Dvijotham, and P. Kohli, “Achieving verified robustness to symbol substitutions via interval bound propagation,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 4083–4093
2019
Cited alongside, same era.
Y. Zhang, S. Sun, M. Galley, Y.-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and W. B. Dolan, “Dialogpt: Large-scale generative pre-training for conversational response generation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , 2020, pp. 270–278
2020
Cited alongside, same era.
K. Shuster, D. Ju, S. Roller, E. Dinan, Y.-L. Boureau, and J. Weston, “The dialogue dodecathlon: Open-domain knowledge and image grounded conversational agents,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 2453–2470
2020
Cited alongside, same era.
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits, “Is bert really robust? a strong baseline for natural language attack on text classification and entailment,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 05, 2020, pp. 8018–8025
2020
Cited alongside, same era.
S. Garg and G. Ramakrishnan, “Bae: Bert-based adversarial examples for text classification,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 6174–6181
2020
Cited alongside, same era.
X. Dong, A. T. Luu, R. Ji, and H. Liu, “Towards robustness against natural language word substitutions,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
D. Zhang, M. Ye, C. Gong, Z. Zhu, and Q. Liu, “Black-box certification with randomized smoothing: A functional optimization based framework,” Advances in Neural Information Processing Systems , vol. 33, pp. 2316–2326, 2020
2020
Cited alongside, same era.
Later among the works it cites.
L. Li, M. Weber, X. Xu, L. Rimanic, B. Kailkhura, T. Xie, C. Zhang, and B. Li, “Tss: Transformation-specific smoothing for robustness certification,” in Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security , 2021, pp. 535–557
2021
Later among the works it cites.
H. Liu, J. Jia, and N. Z. Gong, “Pointguard: Provably robust 3d point cloud classification,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 6186–6195
2021
Later among the works it cites.
M. Moradi and M. Samwald, “Evaluating the robustness of neural language models to input perturbations,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 1558–1570
2021
Later among the works it cites.
B. Wang, C. Xu, S. Wang, Z. Gan, Y. Cheng, J. Gao, A. H. Awadallah, and B. Li, “Adversarial glue: A multi-task benchmark for robustness evaluation of language models,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , 2021
2021
Later among the works it cites.
Y. Yan, R. Li, S. Wang, F. Zhang, W. Wu, and W. Xu, “Consert: A contrastive framework for self-supervised sentence representation transfer,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 5065–5075
2021
Later among the works it cites.
B. Wang, J. Jia, X. Cao, and N. Z. Gong, “Certified robustness of graph neural networks against adversarial structural perturbation,” in Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining , 2021, pp. 1645–1653
2021
Later among the works it cites.
A. Gramatzki, “9 text classification examples in action,” https://levity.ai/blog/9-text-classification-examples , 2022
2022
Later among the works it cites.
D. Lee, S. Moon, J. Lee, and H. O. Song, “Query-efficient and scalable black-box adversarial attacks on discrete sequential data via bayesian optimization,” in International Conference on Machine Learning . PMLR, 2022, pp. 12 478–12 497
2022
Later among the works it cites.
K. Yoo, J. Kim, J. Jang, and N. Kwak, “Detection of adversarial examples in text classification: Benchmark and baseline via robust density estimation,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 3656–3672
2022
Later among the works it cites.
Y. Yang, X. Wang, and K. He, “Robust textual embedding against word-level adversarial attacks,” in Uncertainty in Artificial Intelligence . PMLR, 2022, pp. 2214–2224
2022
Later among the works it cites.
H. Zhao, C. Ma, X. Dong, A. T. Luu, Z.-H. Deng, and H. Zhang, “Certified robustness against natural language attacks by causal intervention,” in International Conference on Machine Learning . PMLR, 2022, pp. 26 958–26 970
2022
Later among the works it cites.
M. Alfarra, A. Bibi, N. Khan, P. H. Torr, and B. Ghanem, “Deformrs: Certifying input deformations with randomized smoothing,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 6, 2022, pp. 6001–6009
2022
Later among the works it cites.
Z. Hao, C. Ying, Y. Dong, H. Su, J. Song, and J. Zhu, “Gsmooth: Certified robustness against semantic transformations via generalized randomized smoothing,” in International Conference on Machine Learning . PMLR, 2022, pp. 8465–8483
2022
Later among the works it cites.
P. Huang, Y. Yang, F. Jia, M. Liu, F. Ma, and J. Zhang, “Word level robustness enhancement: Fight perturbation with perturbation,” in AAAI , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Xie, D. Wang, P.-Y. Chen, J. Xiong, S. Liu, and O. Koyejo, “A word is worth a thousand dollars: Adversarial attack on tweets fools stock prediction,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022, pp. 587–599
2022
Later among the works it cites.
J. C. Pérez, M. Alfarra, S. Giancola, B. Ghanem et al. , “3deformrs: Certifying spatial deformations on point clouds,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 169–15 179
2022
Later among the works it cites.
L. Li, T. Xie, and B. Li, “Sok: Certified robustness for deep neural networks,” in 2023 IEEE symposium on security and privacy (SP) . IEEE, 2023, pp. 1289–1310
2023
Closest in time.
J. Zeng, J. Xu, X. Zheng, and X.-J. Huang, “Certified robustness to text adversarial attacks by randomized [mask],” Computational Linguistics , vol. 49, no. 2, pp. 395–427, 2023
2023
Closest in time.
Z. Huang, N. G. Marchant, K. Lucas, L. Bauer, O. Ohrimenko, and B. Rubinstein, “Rs-del: Edit distance robustness certificates for sequence classifiers via randomized deletion,” Advances in Neural Information Processing Systems , vol. 36, 2023
2023
Closest in time.