Fetching the paper…
Reading the bibliography…
Current large language models (LLMs) provide a strong foundation for large-scale user-oriented natural language tasks.
J. Wei, K. Zou, Eda: Easy data augmentation techniques for boosting performance on text classification tasks (2019) · 1901
Earlier work this paper cites.
E. Wallace, S. Feng, N. Kandpal, M. Gardner, S. Singh, Universal adversarial triggers for attacking and analyzing nlp (2021) · 1908
Earlier work this paper cites.
J. X. Morris, E. Lifland, J. Y. Yoo, J. Grigsby, D. Jin, Y. Qi, Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp (2020) · 2005
Earlier work this paper cites.
T. Shin, Y. Razeghi, R. L. L. I. au2, E. Wallace, S. Singh, Autoprompt: Eliciting knowledge from language models with automatically generated prompts (2020) · 2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
S.-M. Moosavi-Dezfooli, A. Fawzi, P. Frossard, Deepfool: a simple and accurate method to fool deep neural networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, p.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, A. Swami, Practical black-box attacks against machine learning, in: ACCC, 2017, p.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Ebrahimi, A. Rao, D. Lowd, D. Dou, Hotflip: White-box adversarial examples for text classification (2018) · 2018
Earlier work this paper cites.
M. T. Ribeiro, S. Singh, C. Guestrin, Semantically equivalent adversarial rules for debugging NLP models, in: ACL, 2018, p.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Gao, J. Lanchantin, M. L. Soffa, Y. Qi, Black-box generation of adversarial text sequences to evade deep learning classifiers, in: SPW, 2018, p.
2018
Cited alongside, same era.
D. Jin, Z. Jin, J. T. Zhou, P. Szolovits, Is bert really robust? a strong baseline for natural language attack on text classification and entailment, in: AAAI, 2020, p.
2020
Cited alongside, same era.
L. Li, R. Ma, Q. Guo, X. Xue, X. Qiu, BERT-ATTACK: adversarial attack against BERT using BERT, in: EMNLP, 2020, p.
2020
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, A. B. Santoro, S. Chaplot, A. Patra, I. Sutskever, Chatgpt: A language model for conversational agents, OpenAI (2020)
2020
Cited alongside, same era.
OpenAI, Gpt-4 technical report (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, N. Abu-Ghazaleh, Survey of vulnerabilities in large language models revealed by adversarial attacks (2023) · 2023
Later among the works it cites.
S. Goyal, S. Doddapaneni, M. M. Khapra, B. Ravindran, A survey of adversarial defenses and robustness in nlp, ACM Comput. Surv. (2023)
2023
Later among the works it cites.
E. Jones, A. Dragan, A. Raghunathan, J. Steinhardt, Automatically auditing large language models via discrete optimization (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
C. Guo, A. Sablayrolles, H. Jégou, D. Kiela, Gradient-based adversarial attacks against text transformers (2021) · 2021
Cited alongside, same era.
C. L. Canonne, G. Kamath, T. Steinke, The discrete gaussian for differential privacy (2021)
2021
Cited alongside, same era.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, J. Steinhardt, Measuring mathematical problem solving with the math dataset, NeurIPS (2021)
2021
Cited alongside, same era.
F. Perez, I. Ribeiro, Ignore previous prompt: Attack techniques for language models (2022) · 2022
Cited alongside, same era.
E. Wallace, S. Feng, N. Kandpal, M. Gardner, S. Singh, Benchmarking language models’ robustness to semantic perturbations, in: NAACL, 2022, p.
2022
Cited alongside, same era.
Y. Chang, M. Narang, H. Suzuki, G. Cao, J. Gao, Y. Bisk, Webqa: Multihop and multimodal qa, in: CVPR, 2022, pp. 16495–16504
2022
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
Y. Liu, G. Deng, Y. Li, K. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, Y. Liu, Prompt injection attack against llm-integrated applications (2023) · 2023
Later among the works it cites.
2023
Later among the works it cites.
T. Brown, J. Chinchilla, Q. V. Le, B. Mann, A. Roy, D. Saxton, E. Wei, M. Ziegler, Gsm8k: A large-scale dataset for language model training (2023) · 2023
Later among the works it cites.
2023
Later among the works it cites.