Fetching the paper…
Reading the bibliography…
The advent of Large Language Models (LLMs) has marked significant achievements in language processing and reasoning capabilities.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
B. Pang and L. Lee · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher et al · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy et al · 2013
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
X. Chen, C. Liu, et al · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
J. Ebrahimi, A. Rao, et al · 2017
Earlier work this paper cites.
Trojaning attack on neural networks
Y. Liu, S. Ma, et al · 2018
Earlier work this paper cites.
Analyzing federated learning through an adversarial lens
A. N. Bhagoji et al · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
J. Dai, C. Chen, and Y. Li · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
T. Gu, K. Liu, et al · 2019
Earlier work this paper cites.
How to backdoor federated learning
E. Bagdasaryan, A. Veit, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis et al · 2020
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
F. Qi et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin et al · 2020
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
E. Wallace, T. Z. Zhao, S. Feng, and S. Singh · 2020
Earlier work this paper cites.
Dba: Distributed backdoor attacks against federated learning
C. Xie, K. Huang, P. Y. Chen, and B. Li · 2020
Earlier work this paper cites.
Clean-label backdoor attacks on video recognition models
S. Zhao, X. Ma, et al · 2020
Earlier work this paper cites.
Mitigating backdoor attacks in lstm-based text classification systems by backdoor keyword identification
C. Chen and J. Dai · 2021
Earlier work this paper cites.
Badnl: Backdoor attacks against nlp models with semantic-preserving improvements
X. Chen et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe et al · 2021
Earlier work this paper cites.
Triggerless backdoor attack for nlp tasks with clean labels
L. Gan et al · 2021
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
L. Li, D. Song, et al · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
S. Mishra et al · 2021
Cited alongside, same era.
Backdoor pre-trained models can transfer to all
L. Shen et al · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei et al · 2021
Cited alongside, same era.
Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models
W. Yang et al · 2021
Cited alongside, same era.
Backdoor attack against speaker verification
T. Zhai et al · 2021
Cited alongside, same era.
Test-time backdoor mitigation for black-box large language models with defensive demonstrations
W. Mo et al · 2023
Later among the works it cites.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Later among the works it cites.
Hijacking large language models via adversarial in-context learning
Y. Qiang et al · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
J. Rando and F. Tramèr · 2023
Later among the works it cites.
Prompt-specific poisoning attacks on text-to-image generative models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. W. Chung et al · 2022
Cited alongside, same era.
Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages, 2022
J. FitzGerald et al · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli et al · 2022
Cited alongside, same era.
Wedef: Weakly supervised backdoor defense for text classification
L. Jin et al · 2022
Cited alongside, same era.
Backdoor learning: A survey
Y. Li, Y. Jiang, Z. Li, and S.-T. Xia · 2022
Cited alongside, same era.
Holistic evaluation of language models
P. Liang et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang et al · 2022
Cited alongside, same era.
S. Shan, W. Ding, et al · 2023
Later among the works it cites.
On the exploitability of instruction tuning
M. Shu, J. Wang, et al · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model, 2023
R. Taori et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron et al · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
A. Wan et al · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
J. Xu et al · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
J. Yan et al · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Y. Yao et al · 2023
Later among the works it cites.
How do large language models capture the ever-changing world knowledge? a review of recent advances
Z. Zhang et al · 2023
Later among the works it cites.
Lima: Less is more for alignment
C. Zhou et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, et al · 2023
Later among the works it cites.
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al · 2024
Closest in time.
Membership inference attacks against in-context learning
R. Wen, Z. Li, M. Backes, and Y. Zhang · 2024
Closest in time.
Continual learning for large language models: A survey
T. Wu et al · 2024
Closest in time.
Shadowcast: Stealthy data poisoning attacks against vision-language models
Y. Xu et al · 2024
Closest in time.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
S. Zhao, M. Jia, L. A. Tuan, F. Pan, and J. Wen · 2024
Closest in time.
Generative large language model—powered conversational ai app for personalized risk assessment: Case study in covid-19
M. A. Roshani, X. Zhou, Y. Qiang, S. Suresh, S. Hicks, U. Sethuraman, and D. Zhu · 2025
Closest in time.