Fetching the paper…
Reading the bibliography…
Open-source Large Language Models (LLMs) have recently gained popularity because of their comparable performance to proprietary LLMs.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out . Association for Computational Linguistics, 2004, pp. 74–81
2004
Earlier work this paper cites.
2013
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in ICML , 2019
2019
Earlier work this paper cites.
J. Lin, L. Xu, Y. Liu, and X. Zhang, “Composite backdoor attack for deep neural network by mixing existing benign features,” in CCS , 2020
2020
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in ICLR , 2020
2020
Earlier work this paper cites.
S. Li, H. Liu, T. Dong, B. Z. H. Zhao, M. Xue, H. Zhu, and J. Lu, “Hidden backdoors in human-centric language models,” in CCS , 2021
2021
Earlier work this paper cites.
E. Wallace, T. Z. Zhao, S. Feng, and S. Singh, “Concealed data poisoning attacks on NLP models,” in NAACL-HLT , 2021
2021
Earlier work this paper cites.
L. Shen, S. Ji, X. Zhang, J. Li, J. Chen, J. Shi, C. Fang, J. Yin, and T. Wang, “Backdoor pre-trained models can transfer to all,” in CCS , 2021
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in ICLR , 2021
2021
Earlier work this paper cites.
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui, “MAUVE: measuring the gap between neural text and human text using divergence frontiers,” in NeurIPS , 2021
2021
Earlier work this paper cites.
A. Azizi, I. A. Tahmid, A. Waheed, N. Mangaokar, J. Pu, M. Javed, C. K. Reddy, and B. Viswanath, “T-miner: A generative approach to defend against trojan attacks on dnn-based text classification,” in USENIX Security , 2021
2021
Earlier work this paper cites.
H. Jia, M. Yaghini, C. A. Choquette-Choo, N. Dullerud, A. Thudi, V. Chandrasekaran, and N. Papernot, “Proof-of-learning: Definitions and practice,” in IEEE S&P , 2021
2021
Earlier work this paper cites.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in ICLR , 2022
2022
Earlier work this paper cites.
B. Ichter, A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, D. Kalashnikov, S. Levine, Y. Lu, C. Parada, K. Rao, P. Sermanet, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, M. Yan, N. Brown, M. Ahn, O. Cortes, N. Sievers, C. Tan, S. Xu, D. Reyes, J. Rettinghouse, J. Quiambao, P. Pastor, L. Luu, K. Lee, Y. Kuang, S. Jesmonth, N. J. Joshi, K. Jeffrey, R. J. Ruano, J. Hsu, K. Gopalakrishnan, B. David, A. Zeng, and C. K. Fu, “Do as I can, not as I say: Grounding language in robotic affordances,” in Conference on Robot Learning, CoRL 2022 , ser. Proceedings of Machine Learning Research, vol. 205. PMLR, 2022, pp. 287–318
2022
Earlier work this paper cites.
X. Pan, M. Zhang, B. Sheng, J. Zhu, and M. Yang, “Hidden trigger backdoor attack on NLP models via linguistic style manipulation,” in USENIX Security , 2022
2022
Earlier work this paper cites.
H. Chase, “LangChain,” Oct. 2022. [Online]. Available: https://github.com/hwchase17/langchain
2022
Earlier work this paper cites.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in ICLR , 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in NeurIPS , 2022
2022
Earlier work this paper cites.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,” Transactions on Machine Learning Research , 2022, survey Certification
2022
Earlier work this paper cites.
A. Salem, R. Wen, M. Backes, S. Ma, and Y. Zhang, “Dynamic backdoor attacks against machine learning models,” in EuroS&P , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Liu, G. Shen, G. Tao, S. An, S. Ma, and X. Zhang, “Piccolo: Exposing complex backdoors in NLP transformer models,” in IEEE S&P , 2022
2022
Earlier work this paper cites.
S. Li, T. Dong, B. Z. H. Zhao, M. Xue, S. Du, and H. Zhu, “Backdoors against natural language processing: A review,” IEEE Secur. Priv. , vol. 20, no. 5, pp. 50–59, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
“The LLaMA ecosystem: Past, present, and future,” https://ai.meta.com/blog/llama-2-updates-connect-2023/, Sep. 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
A. Shostack, D. Hasse, and R. Kukreja. (2023) Understanding the risks of deploying LLMs in your enterprise. https://www.moveworks.com/insights/risks-of-deploying-llms-in-your-enterprise
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpacaeval: An automatic evaluator of instruction-following models,” https://github.com/tatsu-lab/alpaca_eval , 2023
2023
Closest in time.
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics (TACL) , 2023
2023
Closest in time.
2023
Closest in time.
THUDM, “Chatglm2-6B,” 2023. [Online]. Available: https://github.com/THUDM/ChatGLM2-6B
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
K. Aslett, Z. Sanderson, W. Godel, N. Persily, J. Nagler, and J. A. Tucker, “Online searches to evaluate misinformation can increase its perceived veracity,” Nature , Dec. 2023
2023
Cited alongside, same era.
“Guide: Large language models-generated fraud, malware, and vulnerabilities,” https://fingerprint.com/blog/large-language-models-llm-fraud-malware-guide/, 2023
2023
Cited alongside, same era.
M. Kan. (2023) After wormGPT, fraudGPT emerges to help scammers steal your data. https://www.pcmag.com/news/after-wormgpt-fraudgpt-emerges-to-help-scammers-steal-your-data
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following LLaMA model,” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Cited alongside, same era.
M. Shu, J. Wang, C. Zhu, J. Geiping, C. Xiao, and T. Goldstein, “On the exploitability of instruction tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
K. Mei, Z. Li, Z. Wang, Y. Zhang, and S. Ma, “NOTABLE: transferable backdoor attacks against prompt-based NLP models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023 . Association for Computational Linguistics, 2023, pp. 15 551–15 565
2023
Closest in time.
N. Gu, P. Fu, X. Liu, Z. Liu, Z. Lin, and W. Wang, “A gradient control method for backdoor attacks on parameter-efficient tuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 3508–3520
2023
Closest in time.
2023
Closest in time.
A. Wan, E. Wallace, S. Shen, and D. Klein, “Poisoning language models during instruction tuning,” in International Conference on Machine Learning, ICML 2023 , ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 2023, pp. 35 413–35 425
2023
Closest in time.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “HuggingGPT: Solving AI tasks with chatGPT and its friends in hugging face,” in Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Closest in time.
“AutoGPT,” https://github.com/Significant-Gravitas/Auto-GPT, 2023
2023
Closest in time.
“Babyagi,” https://github.com/yoheinakajima/babyagi, 2023
2023
Closest in time.
T. Dong, S. Li, G. Chen, M. Xue, H. Zhu, and Z. Liu, “RAI2: Responsible identity audit governing the artificial intelligence.” in NDSS , 2023
2023
Closest in time.
Y. Wen and S. Chaudhuri, “Batched low-rank adaptation of foundation models,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
B. Beaumont-Thomas. (2024) Taylor swift deepfake pornography sparks renewed calls for us legislation. https://www.theguardian.com/music/2024/jan/26/taylor-swift-deepfake-pornography-sparks-renewed-calls-for-us-legislation
2024
Closest in time.
“The world’s first bug bounty platform for ai/ml,” https://huntr.com/ , 2024
2024
Closest in time.
L. Sun, Y. Huang, H. Wang, S. Wu, Q. Zhang, C. Gao, Y. Huang, W. Lyu, Y. Zhang, X. Li, Z. Liu, Y. Liu, Y. Wang, Z. Zhang, B. Kailkhura, C. Xiong, C. Xiao, C. Li, E. Xing, F. Huang, H. Liu, H. Ji, H. Wang, H. Zhang, H. Yao, M. Kellis, M. Zitnik, M. Jiang, M. Bansal, J. Zou, J. Pei, J. Liu, J. Gao, J. Han, J. Zhao, J. Tang, J. Wang, J. Mitchell, K. Shu, K. Xu, K.-W. Chang, L. He, L. Huang, M. Backes, N. Z. Gong, P. S. Yu, P.-Y. Chen, Q. Gu, R. Xu, R. Ying, S. Ji, S. Jana, T. Chen, T. Liu, T. Zhou, W. Wang, X. Li, X. Zhang, X. Wang, X. Xie, X. Chen, X. Wang, Y. Liu, Y. Ye, Y. Cao, Y. Chen, and Y. Zhao, “Trustllm: Trustworthiness in large language models,” 2024
2024
Closest in time.
C. Wei, W. Meng, Z. Zhang, M. Chen, M. Zhao, W. Fang, L. Wang, Z. Zhang, and W. Chen, “Lmsanitator: Defending prompt-tuning against task-agnostic backdoors,” in NDSS , 2024
2024
Closest in time.
S. Zhao, L. Gan, L. A. Tuan, J. Fu, L. Lyu, M. Jia, and J. Wen, “Defending against weight-poisoning backdoor attacks for parameter-efficient fine-tuning,” NAACL Finding , 2024
2024
Closest in time.
J. Xu, M. Ma, F. Wang, C. Xiao, and M. Chen, “Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 3111–3126
2024
Closest in time.
2024
Closest in time.
J. Rando and F. Tramèr, “Universal jailbreak backdoors from poisoned human feedback,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
S. Li, X. Wang, M. Xue, H. Zhu, Z. Zhang, Y. Gao, W. Wu, and X. S. Shen, “Yes, one-bit-flip matters! universal dnn model inference depletion with runtime code fault injection,” in Proceedings of the 33th USENIX Security Symposium , 2024
2024
Closest in time.
2024
Closest in time.
S. S. Roy, P. Thota, K. V. Naragam, and S. Nilizadeh, “From chatbots to phishbots?: Phishing scam generation in commercial large language models,” in IEEE Symposium on Security and Privacy (SP) , 2024
2024
Closest in time.
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=hTEGyKf0dZ
2024
Closest in time.