Fetching the paper…
Reading the bibliography…
Open-source large language models (LLMs) have become increasingly popular among both the general public and industry, as they can be customized, fine-tuned, and freely used.
Paraphrasing for Style
Wei Xu, Alan Ritter, Bill Dolan, Ralph Grishman, and Colin Cherry · 2012
Earlier work this paper cites.
Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Badnets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Grag · 2017
Earlier work this paper cites.
Style Transfer from Non-Parallel Text by Cross-Alignment
Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi S. Jaakkola · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Trojaning Attack on Neural Networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2018
Earlier work this paper cites.
A Backdoor Attack Against LSTM-Based Text Classification Systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Reformulating Unsupervised Style Transfer as Paraphrase Generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer · 2020
Earlier work this paper cites.
Weight Poisoning Attacks on Pretrained Models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
HateBERT: Retraining BERT for Abusive Language Detection in English
Tommaso Caselli, Valerio Basile, Jelena Mitrovic, and Michael Granitzer · 2021
Earlier work this paper cites.
BadNL: Backdoor Attacks Against NLP Models with Semantic-preserving Improvements
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
PalmTree: Learning an Assembly Language Model for Instruction Embedding
Xuezixiang Li, Yu Qu, and Heng Yin · 2021
Earlier work this paper cites.
ClipCap: CLIP Prefix for Image Captioning
Ron Mokady, Amir Hertz, and Amit H. Bermano · 2021
Earlier work this paper cites.
ONION: A Simple and Effective Defense Against Textual Backdoor Attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Earlier work this paper cites.
Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He · 2021
Earlier work this paper cites.
Rethinking Stealthiness of Backdoor Attack against NLP Models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun · 2021
Earlier work this paper cites.
Challenges in Automated Debiasing for Toxic Language Detection
Xuhui Zhou, Maarten Sap, Swabha Swayamdipta, Yejin Choi, and Noah A. Smith · 2021
Cited alongside, same era.
Poisoning and Backdooring Contrastive Learning
Nicholas Carlini and Andreas Terzis · 2022
Cited alongside, same era.
PPT: Backdoor Attacks on Pre-trained Models via Poisoned Prompt Tuning
Wei Du, Yichun Zhao, Boqun Li, Gongshen Liu, and Shilin Wang · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Cited alongside, same era.
Watermarking Pre-trained Language Models with Backdooring
Chenxi Gu, Chengsong Huang, Xiaoqing Zheng, Kai-Wei Chang, and Cho-Jui Hsieh · 2022
Cited alongside, same era.
PLMmark: A Secure and Robust Black-Box Watermarking Framework for Pre-trained Language Models
Peixuan Li, Pengzhou Cheng, Fangqi Li, Wei Du, Haodong Zhao, and Gongshen Liu · 2023
Later among the works it cites.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Later among the works it cites.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Later among the works it cites.
Crosslingual Generalization through Multitask Finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinyuan Jia, Yupei Liu, and Neil Zhenqiang Gong · 2022
Cited alongside, same era.
Dynamic Backdoor Attacks Against Machine Learning Models
Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma, and Yang Zhang · 2022
Cited alongside, same era.
Data Poisoning Attacks Against Multimodal Encoders
Ziqing Yang, Xinlei He, Zheng Li, Michael Backes, Mathias Humbert, Pascal Berrang, and Yang Zhang · 2022
Cited alongside, same era.
Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
The Falcon Series of Open Language Models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Cited alongside, same era.
A Test Collection of Synthetic Documents for Training Rankers: ChatGPT vs. Human Experts
Arian Askari, Mohammad Aliannejadi, Evangelos Kanoulas, and Suzan Verberne · 2023
Cited alongside, same era.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung · 2023
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Can Large Language Models Reason about Program Invariants?
Kexin Pei, David Bieber, Kensen Shi, Charles Sutton, and Pengcheng Yin · 2023
Later among the works it cites.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
Prompt Stealing Attacks Against Text-to-Image Generation Models
Xinyue Shen, Yiting Qu, Michael Backes, and Yang Zhang · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2023
Later among the works it cites.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing · 2023
Later among the works it cites.
GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang · 2023
Later among the works it cites.
HuRef: HUman-REadable Fingerprint for Large Language Models
Boyi Zeng, Chenghu Zhou, Xinbing Wang, and Zhouhan Lin · 2023
Later among the works it cites.
Effective Prompt Extraction from Language Models
Yiming Zhang, Nicholas Carlini, and Daphne Ippolito · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Benchmarking Large Language Models in Retrieval-Augmented Generation
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun · 2024
Closest in time.
Prompt Stealing Attacks Against Large Language Models
Zeyang Sha and Yang Zhang · 2024
Closest in time.
Instructional Fingerprinting of Large Language Models
Jiashu Xu, Fei Wang, Mingyu Derek Ma, Pang Wei Koh, Chaowei Xiao, and Muhao Chen · 2024
Closest in time.
PRSA: Prompt Reverse Stealing Attacks against Large Language Models
Yong Yang, Xuhong Zhang, Yi Jiang, Xi Chen, Haoyu Wang, Shouling Ji, and Zonghui Wang · 2024
Closest in time.