Fetching the paper…
Reading the bibliography…
Large language models play a crucial role in modern natural language processing technologies.
Discovering Language Model Behaviors with Model-Written Evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Earlier work this paper cites.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Earlier work this paper cites.
Whose Opinions Do Language Models Reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Earlier work this paper cites.
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations, October 2023
Zeming Wei, Yifei Wang, and Yisen Wang · 2023
Earlier work this paper cites.
Multi-step Jailbreaking Privacy Attacks on ChatGPT, November 2023
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song · 2023
Earlier work this paper cites.
Defending ChatGPT against Jailbreak Attack via Self-Reminder, June 2023
Fangzhao Wu, Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, and Xing Xie · 2023
Earlier work this paper cites.
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM, December 2023
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen · 2023
Cited alongside, same era.
Low-Resource Languages Jailbreak GPT-4, October 2023
Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks, November 2023
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas · 2023
Later among the works it cites.
Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study, March 2024
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, Kailong Wang, and Yang Liu · 2024
Closest in time.
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning, March 2024
Shuai Zhao, Leilei Gan, Luu Anh Tuan, Jie Fu, Lingjuan Lyu, Meihuizi Jia, and Jinming Wen · 2024
Closest in time.
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework, March 2024
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, and Mei Si · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, T. Strzalkowski, and Mei Si · 2023
Cited alongside, same era.
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju · 2024
Closest in time.