Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of generating obstinate (over-stability) adversarial examples by word substitution in NLP, where input text is meaningfully changed but the model's prediction does not, even though it should.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
WordNet: A Lexical Database for English
Miller, G. A. 1995 · 1995
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
Word Shape Matters: Robust Machine Translation with Visual Embedding
Wang, H.; Zhang, P.; and Xing, E. P. 2020 · 2010
Earlier work this paper cites.
Learning Word Vectors for Sentiment Analysis
Maas, A. L.; Daly, R. E.; Pham, P. T.; Huang, D.; Ng, A. Y.; and Potts, C. 2011 · 2011
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I. J.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Deep Reinforcement Learning for Dialogue Generation
Li, J.; Monroe, W.; Ritter, A.; Jurafsky, D.; Galley, M.; and Gao, J. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues
Serban, I. V.; Sordoni, A.; Lowe, R.; Charlin, L.; Pineau, J.; Courville, A. C.; and Bengio, Y. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Earlier work this paper cites.
Generating Natural Language Adversarial Examples
Alzantot, M.; Sharma, Y.; Elgohary, A.; Ho, B.-J.; Srivastava, M.; and Chang, K.-W. 2018 · 2018
Earlier work this paper cites.
Decoupled Classifiers for Group-Fair and Efficient Machine Learning
Dwork, C.; Immorlica, N.; Kalai, A. T.; and Leiserson, M. 2018 · 2018
Earlier work this paper cites.
On Adversarial Examples for Character-Level Neural Machine Translation
Ebrahimi, J.; Lowd, D.; and Dou, D. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-Box Adversarial Examples for Text Classification
Ebrahimi, J.; Rao, A.; Lowd, D.; and Dou, D. 2018 · 2018
Cited alongside, same era.
Pathologies of Neural Models Make Interpretations Difficult
Feng, S.; Wallace, E.; Grissom II, A.; Iyyer, M.; Rodriguez, P.; and Boyd-Graber, J. 2018 · 2018
Cited alongside, same era.
Adversarial Over-Sensitivity and Over-Stability Strategies for Dialogue Models
Niu, T.; and Bansal, M. 2018 · 2018
Cited alongside, same era.
Semantically Equivalent Adversarial Rules for Debugging NLP models
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Excessive invariance causes adversarial vulnerability
Jacobsen, J.-H.; Behrmann, J.; Zemel, R. S.; and Bethge, M. 2019 · 2019
BERT-ATTACK: Adversarial Attack Against BERT Using BERT
Li, L.; Ma, R.; Guo, Q.; Xue, X.; and Qiu, X. 2020 · 2020
Later among the works it cites.
Reevaluating Adversarial Examples in Natural Language
Morris, J.; Lifland, E.; Lanchantin, J.; Ji, Y.; and Qi, Y. 2020 · 2020
Later among the works it cites.
Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
Tramer, F.; Behrmann, J.; Carlini, N.; Papernot, N.; and Jacobsen, J.-H. 2020 · 2020
Later among the works it cites.
Towards Verified Robustness under Text Deletion Interventions
Welbl, J.; Huang, P.; Stanforth, R.; Gowal, S.; Dvijotham, K. D.; Szummer, M.; and Kohli, P. 2020a · 2020
Later among the works it cites.
Undersensitivity in Neural Reading Comprehension
Welbl, J.; Minervini, P.; Bartolo, M.; Stenetorp, P.; and Riedel, S. 2020b · 2020
Later among the works it cites.
Adversarial Attacks on Deep-learning Models in Natural Language Processing: A Survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
McCoy, T.; Pavlick, E.; and Linzen, T. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
Generating Natural Language Adversarial Examples through Probability Weighted Word Saliency
Ren, S.; Deng, Y.; He, K.; and Che, W. 2019 · 2019
Cited alongside, same era.
Generalization to Mitigate Synonym Substitution Attacks
Alshemali, B.; and Kalita, J. 2020 · 2020
Cited alongside, same era.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2020
Cited alongside, same era.
Seq2Sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples
Cheng, M.; Yi, J.; Chen, P.-Y.; Zhang, H.; and Hsieh, C.-J. 2020 · 2020
Cited alongside, same era.
Zhang, W. E.; Sheng, Q. Z.; Alhazmi, A.; and Li, C. 2020 · 2020
Later among the works it cites.
To what extent do human explanations of model behavior align with actual model behavior?
Prasad, G.; Nie, Y.; Bansal, M.; Jia, R.; Kiela, D.; and Williams, A. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
Semantics Altering Modifications for Evaluating Comprehension in Machine Reading
Schlegel, V.; Nenadic, G.; and Batista-Navarro, R. 2021 · 2021
Later among the works it cites.
Neural Natural Logic Inference for Interpretable Question Answering
Shi, J.; Ding, X.; Du, L.; Liu, T.; and Qin, B. 2021 · 2021
Later among the works it cites.
Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble
Zhou, Y.; Zheng, X.; Hsieh, C.-J.; Chang, K.-W.; and Huang, X. 2021 · 2021
Later among the works it cites.
Balanced Adversarial Training: Balancing Tradeoffs between Fickleness and Obstinacy in NLP Models
Chen, H.; Ji, Y.; and Evans, D. 2022 · 2022
Later among the works it cites.
Discovering the hidden vocabulary of DALLE-2
Daras, G.; and Dimakis, A. G. 2022 · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
Measure and Improve Robustness in NLP Models: A Survey
Wang, X.; Wang, H.; and Yang, D. 2022 · 2022
Later among the works it cites.
CHATGPT: Optimizing language models for dialogue
OpenAI. 2023 · 2023
Closest in time.
Adversarial Examples for Evaluating Reading Comprehension Systems
Jia, R.; and Liang, P. 2017 · 2031
Closest in time.
Modeling Disclosive Transparency in NLP Application Descriptions
Saxon, M.; Levy, S.; Wang, X.; Albalak, A.; and Wang, W. Y. 2021 · 2037
Closest in time.