Fetching the paper…
Reading the bibliography…
Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels.
Binary codes capable of correcting deletions, insertions, and reversals
Levenshtein, V. I. et al · 1966
Earlier work this paper cites.
Validation of subgradient optimization
Held, M., Wolfe, P., and Crowder, H. P · 1974
Earlier work this paper cites.
The Significance of Letter Position in Word Recognition
Rawlinson, G · 1976
Earlier work this paper cites.
Tokenization as the initial phase in NLP
Webster, J. J. and Kit, C · 1992
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
Tokenisation and sentence segmentation
Palmer, D. D · 2000
Earlier work this paper cites.
Psycholinguistic evidence on scrambled letters in reading, 2003
Davis, M · 2003
Earlier work this paper cites.
Fast approximate search in large dictionaries
Mihov, S. and Schulz, K. U · 2004
Earlier work this paper cites.
Ag’s corpus of news articles, 2005
Gulli, A · 2005
Earlier work this paper cites.
Universal levenshtein automata. building and properties
Mitankin, P. N · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
On the levenshtein automaton and the size of the neighbourhood of a word
Touzet, H · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., and Chang, K.-W · 2018
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Belinkov, Y. and Bisk, Y · 2018
Cited alongside, same era.
Universal sentence encoder for English
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., St. John, R., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., Strope, B., and Kurzweil, R · 2018
Cited alongside, same era.
HotFlip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2018
Cited alongside, same era.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Gao, J., Lanchantin, J., Soffa, M. L., and Qi, Y · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Robust encodings: A framework for combating adversarial typos
Jones, E., Jia, R., Raghunathan, A., and Liang, P · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Later among the works it cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Li, L., Ma, R., Guo, Q., Xue, X., and Qiu, X · 2020
Later among the works it cites.
Reevaluating adversarial examples in natural language
Morris, J., Lifland, E., Lanchantin, J., Ji, Y., and Qi, Y · 2020
Later among the works it cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
Morris, J., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., and Qi, Y · 2020
Later among the works it cites.
Imitation attacks and defenses for black-box machine translation systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kudo, T. and Richardson, J · 2018
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
Why do adversarial attacks transfer? explaining transferability of evasion and poisoning attacks
Demontis, A., Melis, M., Pintor, M., Jagielski, M., Biggio, B., Oprea, A., Nita-Rotaru, C., and Roli, F · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Discrete adversarial attacks and submodular optimization with applications to text classification
Lei, Q., Wu, L., Chen, P.-Y., Dimakis, A., Dhillon, I. S., and Witbrock, M. J · 2019
Cited alongside, same era.
Wallace, E., Stern, M., and Song, D · 2020
Later among the works it cites.
Greedy attack and gumbel attack: Generating adversarial examples for discrete data
Yang, P., Chen, J., Hsieh, C.-J., Wang, J.-L., and Jordan, M. I · 2020
Later among the works it cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Later among the works it cites.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Later among the works it cites.
Fast WordPiece tokenization
Song, X., Salcianu, A., Song, Y., Dopson, D., and Zhou, D · 2021
Later among the works it cites.
Query-efficient and scalable black-box adversarial attacks on discrete sequential data via bayesian optimization
Lee, D., Moon, S., Lee, J., and Song, H. O · 2022
Later among the works it cites.
Character-level white-box adversarial attacks against transformers via attachable subwords substitution
Liu, A., Yu, H., Hu, X., Li, S., Lin, L., Ma, F., Yang, Y., and Wen, L · 2022
Later among the works it cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W., Ippolito, D., Tramèr, F., and Schmidt, L · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Later among the works it cites.
How do humans perceive adversarial text? a reality check on the validity and naturalness of word-based adversarial attacks
Dyrmishi, S., Ghamizi, S., and Cordy, M · 2023
Later among the works it cites.
Textgrad: Advancing robustness evaluation in NLP by gradient-driven optimization
Hou, B., Jia, J., Zhang, Y., Zhang, G., Zhang, Y., Liu, S., and Chang, S · 2023
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models, 2023
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Targeted adversarial attacks against neural machine translation
Sadrizadeh, S., Aghdam, A. D., Dolamic, L., and Frossard, P · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.