Fetching the paper…
Reading the bibliography…
The adoption of large language models (LLMs) in many applications, from customer service chat bots and software development assistants to more capable agentic systems necessitates research into how to secure these systems.
Bad characters: Imperceptible nlp attacks
Boucher, N.; Shumailov, I.; Anderson, R.; and Papernot, N. 2022 · 2004
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering
Karpukhin, V.; Oğuz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; and tau Yih, W. 2020 · 2004
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I. J.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
CVE-2019-20634
Pearce, W.; and Landers, N. 2019 · 2019
Earlier work this paper cites.
Text embeddings by weakly-supervised contrastive pre-training
Wang, L.; Yang, N.; Huang, X.; Jiao, B.; Yang, L.; Jiang, D.; Majumder, R.; and Wei, F. 2022 · 2022
Earlier work this paper cites.
Increasing confidence in adversarial robustness evaluations
Zimmermann, R. S.; Brendel, W.; Tramer, F.; and Carlini, N. 2022 · 2022
Earlier work this paper cites.
Summon a demon and bind it: A grounded theory of llm red teaming in the wild
Inie, N.; Stray, J.; and Derczynski, L. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Cited alongside, same era.
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Lin, Z.; Wang, Z.; Tong, Y.; Wang, Y.; Guo, Y.; Wang, Y.; and Shang, J. 2023 · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A.; Zampetakis, M.; Kassianik, P.; Nelson, B.; Anderson, H.; Singer, Y.; and Karbasi, A. 2023 · 2023
Cited alongside, same era.
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Lee, C.; Roy, R.; Xu, M.; Raiman, J.; Shoeybi, M.; Catanzaro, B.; and Ping, W. 2024 · 2024
Closest in time.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2024 · 2024
Closest in time.
Embedding And Clustering Your Data Can Improve Contrastive Pretraining
Merrick, L. 2024 · 2024
Closest in time.
Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024 · 2024
Closest in time.
Defending llms against jailbreaking attacks via backtranslation
Wang, Y.; Shi, Z.; Bai, A.; and Hsieh, C.-J. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rebedea, T.; Dinu, R.; Sreedhar, M.; Parisien, C.; and Cohen, J. 2023 · 2023
Cited alongside, same era.
Fundamental limitations of alignment in large language models
Wolf, Y.; Wies, N.; Avnery, O.; Levine, Y.; and Shashua, A. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Cited alongside, same era.
garak: A Framework for Security Probing Large Language Models
Derczynski, L.; Galinkin, E.; Martin, J.; Majumdar, S.; and Inie, N. 2024 · 2024
Cited alongside, same era.
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2024 · 2024
Closest in time.
Improving alignment and robustness with circuit breakers
Zou, A.; Phan, L.; Wang, J.; Duenas, D.; Lin, M.; Andriushchenko, M.; Kolter, J. Z.; Fredrikson, M.; and Hendrycks, D. 2024 · 2024
Closest in time.