Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated significant utility in a wide range of applications; however, their deployment is plagued by security vulnerabilities, notably jailbreak attacks.
1908
Earlier work this paper cites.
S. Bird, E. Klein, and E. Loper, Natural language processing with Python: analyzing text with the natural language toolkit . " O’Reilly Media, Inc.", 2009
2009
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Earlier work this paper cites.
Q. Gu, “Llm-based code generation method for golang compiler testing,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2023, pp. 2201–2203
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Rajani, L. Tunstall, E. Beeching, N. Lambert, A. M. Rush, and T. Wolf, “No robots,” https://huggingface.co/datasets/HuggingFaceH4/no_robots , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” 2023
2023
Earlier work this paper cites.
C. Barrett, B. Boyd, E. Bursztein, N. Carlini, B. Chen, J. Choi, A. R. Chowdhury, M. Christodorescu, A. Datta, S. Feizi et al. , “Identifying and mitigating the security risks of generative ai,” Foundations and Trends® in Privacy and Security , vol. 6, no. 1, pp. 1–52, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
“Anthropic \ introducing 100k context windows,” https://www.anthropic.com/index/100k-context-windows , (Accessed on 10/31/2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Cheng, M. Georgopoulos, V. Cevher, and G. Chrysos, “Leveraging context in jailbreaking attacks,” in ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
X. Lin, W. Wang, Y. Li, S. Yang, F. Feng, Y. Wei, and T.-S. Chua, “Data-efficient fine-tuning for llm-based recommendation,” in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2024, pp. 365–374
2024
Cited alongside, same era.
S. Dai, C. Xu, S. Xu, L. Pang, Z. Dong, and J. Xu, “Bias and unfairness in information retrieval systems: New challenges in the llm era,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 6437–6447
2024
Cited alongside, same era.
W. Xu, G. Zhu, X. Zhao, L. Pan, L. Li, and W. Wang, “Pride and prejudice: Llm amplifies self-bias in self-refinement,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 15 474–15 492
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
“Hugging face datasets,” https://huggingface.co/datasets , 2024, accessed: 2024-11-09
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu, “Masterkey: Automated jailbreaking of large language model chatbots,” in Proc. ISOC NDSS , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Team, “Meta llama guard 2,” https://github.com/meta-llama/PurpleLlama/blob/main/Llama-Guard2/MODEL_CARD.md , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Fu, Y. Li, W. Xiao, C. Liu, and Y. Dong, “Safety alignment in NLP tasks: Weakly aligned summarization as an in-context attack,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 8483–8502. [Online]. Available: https://aclanthology.org/2024.acl-long.461
2024
Later among the works it cites.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
AI@Meta, “Llama 3 model card,” 2024. [Online]. Available: https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
2024
Later among the works it cites.
“Openai moderation api,” https://platform.openai.com/docs/guides/moderation , 2024, accessed: 2024-11-10
2024
Later among the works it cites.
C. Zheng, F. Yin, H. Zhou, F. Meng, J. Zhou, K.-W. Chang, M. Huang, and N. Peng, “On prompt-driven safeguarding for large language models,” in Forty-first International Conference on Machine Learning , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Xu, Y. Liu, G. Deng, Y. Li, and S. Picek, “A comprehensive study of jailbreak attack versus defense for large language models,” in Findings of the Association for Computational Linguistics ACL 2024 , 2024, pp. 7432–7449
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Yu, X. Lin, Z. Yu, and X. Xing, “ { \{ LLM-Fuzzer } \} : Scaling assessment of large language model jailbreaks,” in 33rd USENIX Security Symposium (USENIX Security 24) , 2024, pp. 4657–4674
2024
Later among the works it cites.