Fetching the paper…
Reading the bibliography…
A novel hack involving Large Language Models (LLMs) has emerged, exploiting adversarial suffixes to deceive models into generating perilous responses.
Monitor alarm fatigue: an integrative review
Maria Cvach · 2012
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples, 2016
Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Dp-gan: Diversity-promoting generative adversarial network for generating informative and diversified text, 2018
Jingjing Xu, Xuancheng Ren, Junyang Lin, and Xu Sun · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Earlier work this paper cites.
DocRED: A large-scale document-level relation extraction dataset
Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, and Maosong Sun · 2019
Earlier work this paper cites.
Adversarial attack and defense of structured prediction models
Wenjuan Han, Liwen Zhang, Yong Jiang, and Kewei Tu · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts, 2020
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV au2, Eric Wallace, and Sameer Singh · 2020
Cited alongside, same era.
Reclor: A reading comprehension dataset requiring logical reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng · 2020
Cited alongside, same era.
Globally-robust neural networks, 2021
Klas Leino, Zifan Wang, and Matt Fredrikson · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgments, 2022
Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving · 2022
Open sesame! universal black box jailbreaking of large language models, 2023
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Closest in time.
Platypus: Quick, cheap, and powerful refinement of llms, 2023
Ariel N. Lee, Cole J. Hunter, and Nataniel Ruiz · 2023
Closest in time.
Rain: Your language models can align themselves without finetuning, 2023
Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang · 2023
Closest in time.
Tapir: Trigger action platform for information retrieval
Annunziata Elefante Mattia Limone, Gaetano Cimino · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Arb: Advanced reasoning benchmark for large language models, 2023
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, and Aran Komatsuzaki · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unsolved problems in ml safety, 2022
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Theoremqa: A theorem-driven question answering dataset
Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia · 2023
Cited alongside, same era.
perplexity, howpublished = https://huggingface.co/docs/transformers/perplexity , note = Accessed: 2023-08-26, 2023
Huggingface · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Jaramilo:gpt4jailbreak, howpublished = https://huggingface.co/datasets/rubend18/chatgpt-jailbreak-prompts , 2023
Rubén Darío Jaramillo · 2023
Cited alongside, same era.
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, and Nikolay Bashlykov · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions, 2023
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Jailbroken: How does llm safety training fail?, 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Fundamental limitations of alignment in large language models, 2023
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher, 2023
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.