Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), which bridge the gap between human language understanding and complex problem-solving, achieve state-of-the-art performance on several NLP tasks, particularly in few-shot and zero-shot settings.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
A rule-based style and grammar checker
Daniel Naber et al · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Biographies, bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira · 2007
Earlier work this paper cites.
Open-sourced dataset protection via backdoor watermarking
Yiming Li, Ziqi Zhang, Jiawang Bai, Baoyuan Wu, Yong Jiang, and Shu-Tao Xia · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation (wmt16)
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, et al · 2016
Earlier work this paper cites.
The iwslt 2016 evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, Rolando Cattoni, and Marcello Federico · 2016
Earlier work this paper cites.
Tag correlation and user social relation based microblog recommendation
Huifang Ma, Meihuizi Jia, Xianghong Lin, and Fuzhen Zhuang · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Earlier work this paper cites.
Hate speech dataset from a white supremacy forum
Ona De Gibert, Naiara Perez, Aitor Garcıa-Pablos, and Montse Cuadros · 2018
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata · 2018
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, Jeffrey Dean, and Sanjay Ghemawat · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
A qualitative comparison of coqa, squad 2.0 and quac
Mark Yatskar · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar · 2019
Earlier work this paper cites.
Can adversarial weight perturbations inject neural backdoors
Siddhant Garg, Adarsh Kumar, Vibhor Goel, and Yingyu Liang · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Cc-news-en: A large english news corpus
Joel Mackenzie, Rodger Benham, Matthias Petri, Johanne R Trippas, J Shane Culpepper, and Alistair Moffat · 2020
Earlier work this paper cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni · 2020
Earlier work this paper cites.
Short text topic modeling with topic distribution quantization and negative sampling decoder
Xiaobao Wu, Chunping Li, Yan Zhu, and Yishu Miao · 2020
Earlier work this paper cites.
T-miner: A generative approach to defend against trojan attacks on dnn-based text classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K Reddy, and Bimal Viswanath · 2021
Earlier work this paper cites.
How should pre-trained language models be fine-tuned towards adversarial robustness?
Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang · 2021
Earlier work this paper cites.
Text backdoor detection using an interpretable rnn abstract model
Ming Fan, Ziliang Si, Xiaofei Xie, Yang Liu, and Ting Liu · 2021
Earlier work this paper cites.
Spectre: Defending against backdoor attacks using robust statistics
Jonathan Hayase, Weihao Kong, Raghav Somani, and Sewoong Oh · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu · 2021
Earlier work this paper cites.
Hidden backdoors in human-centric language models
Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu · 2021
Earlier work this paper cites.
Bfclass: A backdoor-free text classification framework
Zichao Li, Dheeraj Mekala, Chengyu Dong, and Jingbo Shang · 2021
Earlier work this paper cites.
Enriching and controlling global semantics for text summarization
Thong Nguyen, Anh Tuan Luu, Truc Lu, and Tho Quan · 2021
Cited alongside, same era.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Cited alongside, same era.
Bddr: An effective defense against textual backdoor attacks
Kun Shao, Junan Yang, Yang Ai, Hui Liu, and Yu Zhang · 2021
Cited alongside, same era.
Backdoor pre-trained models can transfer to all
Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang · 2021
Cited alongside, same era.
Demon in the variant: Statistical analysis of { \{ DNNs } \} for robust backdoor contamination detection
Di Tang, XiaoFeng Wang, Haixu Tang, and Kehuan Zhang · 2021
Cited alongside, same era.
Concealed data poisoning attacks on nlp models
Haoran Wang and Kai Shu · 2023
Later among the works it cites.
Lmsanitator: Defending prompt-tuning against task-agnostic backdoors
Chengkun Wei, Wenlong Meng, Zhikun Zhang, Min Chen, Minghu Zhao, Wenjing Fang, Lei Wang, Zihui Zhang, and Wenzhi Chen · 2023
Later among the works it cites.
A unified detection framework for inference-stage backdoor defenses
Xun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu, Xuan Bi, Mingyi Hong, and Jie Ding · 2023
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh · 2021
Cited alongside, same era.
Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun · 2021
Cited alongside, same era.
Trojaning language models for fun and profit
Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang · 2021
Cited alongside, same era.
Spinning language models: Risks of propaganda-as-a-service and countermeasures
Eugene Bagdasaryan and Vitaly Shmatikov · 2022
Cited alongside, same era.
Ppt: Backdoor attacks on pre-trained models via poisoned prompt tuning
Wei Du, Yichun Zhao, Boqun Li, Gongshen Liu, and Shilin Wang · 2022
Cited alongside, same era.
Triggerless backdoor attack for nlp tasks with clean labels
Leilei Gan, Jiwei Li, Tianwei Zhang, Xiaoya Li, Yuxian Meng, Fei Wu, Yi Yang, Shangwei Guo, and Chun Fan · 2022
Cited alongside, same era.
Wedef: Weakly supervised backdoor defense for text classification
Lesheng Jin, Zihan Wang, and Jingbo Shang · 2022
Cited alongside, same era.
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2023
Later among the works it cites.
Large language models are better adversaries: Exploring generative clean-label backdoor attacks against text classifiers
Wencong You, Zayd Hammoudeh, and Daniel Lowd · 2023
Later among the works it cites.
Ncl: Textual backdoor defense using noise-augmented contrastive learning
Shengfang Zhai, Qingni Shen, Xiaoyi Chen, Weilong Wang, Cong Li, Yuejian Fang, and Zhonghai Wu · 2023
Later among the works it cites.
Prompting large language model for machine translation: A case study
Biao Zhang, Barry Haddow, and Alexandra Birch · 2023
Later among the works it cites.
Prompt as triggers for backdoor attack: Examining the vulnerability in language models
Shuai Zhao, Jinming Wen, Anh Luu, Junbo Zhao, and Jie Fu · 2023
Later among the works it cites.
Backdoor attacks with input-unique triggers in nlp
Xukun Zhou, Jiwei Li, Tianwei Zhang, Lingjuan Lyu, Muqiao Yang, and Jun He · 2023
Later among the works it cites.
Yulin Chen, Haoran Li, Zihao Zheng, and Yangqiu Song · 2024
Closest in time.
The philosopher’s stone: Trojaning plugins of large language models
Tian Dong, Minhui Xue, Guoxing Chen, Rayne Holland, Shaofeng Li, Yan Meng, Zhen Liu, and Haojin Zhu · 2024
Closest in time.
Light-peft: Lightening parameter-efficient fine-tuning via early pruning
Naibin Gu, Peng Fu, Xiyu Liu, Bowen Shen, Zheng Lin, and Weiping Wang · 2024
Closest in time.
Cbas: Character-level backdoor attacks against chinese pre-trained language models
Xinyu He, Fengrui Hao, Tianlong Gu, and Liang Chang · 2024
Closest in time.
Vaccine: Perturbation-aware alignment for large language model
Tiansheng Huang, Sihao Hu, and Ling Liu · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
Exploring backdoor attacks against large language model-based decision making
Ruochen Jiao, Shaoyuan Xie, Justin Yue, Takami Sato, Lixu Wang, Yixuan Wang, Qi Alfred Chen, and Qi Zhu · 2024
Closest in time.
Revisiting backdoor attacks against large vision-language models
Siyuan Liang, Jiawei Liang, Tianyu Pang, Chao Du, Aishan Liu, Ee-Chien Chang, and Xiaochun Cao · 2024
Closest in time.
Backdoor attacks on dense passage retrievers for disseminating misinformation
Quanyu Long, Yue Deng, LeiLei Gan, Wenya Wang, and Sinno Jialin Pan · 2024
Closest in time.
Cross-context backdoor attacks against graph prompt learning
Xiaoting Lyu, Yufei Han, Wei Wang, Hangwei Qian, Ivor Tsang, and Xiangliang Zhang · 2024
Closest in time.
Backdoor attacks to deep neural networks: A survey of the literature, challenges, and future research directions
Orson Mengara, Anderson Avila, and Tiago H Falk · 2024
Closest in time.
Codepurify: Defend backdoor attacks on neural code models via entropy-based purification
Fangwen Mu, Junjie Wang, Zhuohao Yu, Lin Shi, Song Wang, Mingyang Li, and Qing Wang · 2024
Closest in time.
Backdoor attacks and defenses in federated learning: Survey, challenges and future research directions
Thuy Dung Nguyen, Tuan Nguyen, Phi Le Nguyen, Hieu H Pham, Khoa D Doan, and Kok-Seng Wong · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin · 2024
Closest in time.
Are LLMs good zero-shot fallacy classifiers?
Fengjun Pan, Xiaobao Wu, Zongrui Li, and Anh Tuan Luu · 2024
Closest in time.
Learning to poison large language models during instruction tuning
Yao Qiang, Xiangyu Zhou, Saleh Zare Zade, Mohammad Amin Roshani, Douglas Zytko, and Dongxiao Zhu · 2024
Closest in time.
Competition report: Finding universal jailbreak backdoors in aligned llms
Javier Rando, Francesco Croce, Kryštof Mitka, Stepan Shabalin, Maksym Andriushchenko, Nicolas Flammarion, and Florian Tramèr · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al · 2024
Closest in time.
Dmgnn: Detecting and mitigating backdoor attacks in graph neural networks
Hao Sui, Bing Chen, Jiale Zhang, Chengcheng Zhu, Di Wu, Qinghua Lu, and Guodong Long · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, et al · 2024
Closest in time.
Bdmmt: Backdoor sample detection for language models through model mutation testing
Jiali Wei, Ming Fan, Wenjing Jiao, Wuxia Jin, and Ting Liu · 2024
Closest in time.
AKEW: Assessing knowledge editing in the wild
Xiaobao Wu, Liangming Pan, William Yang Wang, and Anh Tuan Luu · 2024
Closest in time.
Defending pre-trained language models as few-shot learners against backdoor attacks
Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouling Ji, Jinghui Chen, Fenglong Ma, and Ting Wang · 2024
Closest in time.
Nlpsweep: A comprehensive defense scheme for mitigating nlp backdoor attacks
Tao Xiang, Fei Ouyang, Di Zhang, Chunlong Xie, and Hao Wang · 2024
Closest in time.
Atlantis: Aesthetic-oriented multiple granularities fusion network for joint multimodal aspect-based sentiment analysis
Luwei Xiao, Xingjiao Wu, Junjie Xu, Weijie Li, Cheng Jin, and Liang He · 2024
Closest in time.
Trojllm: A black-box trojan prompt attack on large language models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Bölöni, and Qian Lou · 2024
Closest in time.
Watch out for your agents! investigating backdoor threats to llm-based agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun · 2024
Closest in time.
Poisonprompt: Backdoor attack on prompt-based large language models
Hongwei Yao, Jian Lou, and Zhan Qin · 2024
Closest in time.
E-sage: Explainability-based defense against backdoor attacks on graph neural networks
Dingqiang Yuan, Xiaohua Xu, Lei Yu, Tongchang Han, Rongchang Li, and Meng Han · 2024
Closest in time.
Beear: Embedding-based adversarial removal of safety backdoors in instruction-tuned language models
Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song, Bo Li, and Ruoxi Jia · 2024
Closest in time.
Rapid adoption, hidden risks: The dual impact of large language model customization
Rui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang, Yuan Zhang, Michael Backes, Yun Shen, and Yang Zhang · 2024
Closest in time.
Defending against weight-poisoning backdoor attacks for parameter-efficient fine-tuning
Shuai Zhao, Leilei Gan, Luu Anh Tuan, Jie Fu, Lingjuan Lyu, Meihuizi Jia, and Jinming Wen · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Closest in time.
Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in nlp models
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He · 2058
Closest in time.