Fetching the paper…
Reading the bibliography…
LLM have achieved success in many fields but still troubled by problematic content in the training corpora.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Corrigendum to ”The EU Proposal for a General Data Protection Regulation and the roots of the ’right to be forgotten’” [2013] 29 CLSR 229-235
Mantelero, A. 2013 · 2013
Earlier work this paper cites.
Sparks of Artificial General Intelligence: Early experiments with GPT-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M. T.; and Zhang, Y. 2023 · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Carlini, N.; Nasr, M.; Choquette-Choo, C. A.; Jagielski, M.; Gao, I.; Koh, P. W.; Ippolito, D.; Tramèr, F.; and Schmidt, L. 2023 · 2023
Earlier work this paper cites.
Unlearn What You Want to Forget: Efficient Unlearning for LLMs
Chen, J.; and Yang, D. 2023 · 2023
Earlier work this paper cites.
Who’s Harry Potter? Approximate Unlearning in LLMs
Eldan, R.; and Russinovich, M. 2023 · 2023
Earlier work this paper cites.
Knowledge Unlearning for Mitigating Privacy Risks in Language Models
Jang, J.; Yoon, D.; Yang, S.; Cha, S.; Lee, M.; Logeswaran, L.; and Seo, M. 2023 · 2023
Earlier work this paper cites.
Model Sparsity Can Simplify Machine Unlearning
Jia, J.; Liu, J.; Ram, P.; Yao, Y.; Liu, G.; Liu, Y.; Sharma, P.; and Liu, S. 2023 · 2023
Earlier work this paper cites.
Automatically Auditing Large Language Models via Discrete Optimization
Jones, E.; Dragan, A. D.; Raghunathan, A.; and Steinhardt, J. 2023 · 2023
Earlier work this paper cites.
Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Lapid, R.; Langberg, R.; and Sipper, M. 2023 · 2023
Earlier work this paper cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Cited alongside, same era.
Jailbroken: How Does LLM Safety Training Fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023 · 2023
Cited alongside, same era.
DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models
Wu, X.; Li, J.; Xu, M.; Dong, W.; Wu, S.; Bian, C.; and Xiong, D. 2023 · 2023
Cited alongside, same era.
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Casper, S.; Schulze, L.; Patel, O.; and Hadfield-Menell, D. 2024 · 2024
TOFU: A Task of Fictitious Unlearning for LLMs
Maini, P.; Feng, Z.; Schwarzschild, A.; Lipton, Z. C.; and Kolter, J. Z. 2024 · 2024
Closest in time.
Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
Patil, V.; Hase, P.; and Bansal, M. 2024 · 2024
Closest in time.
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Paulus, A.; Zharmagambetov, A.; Guo, C.; Amos, B.; and Tian, Y. 2024 · 2024
Closest in time.
MUSE: Machine Unlearning Six-Way Evaluation for Language Models
Shi, W.; Lee, J.; Huang, Y.; Malladi, S.; Zhao, J.; Holtzman, A.; Liu, D.; Zettlemoyer, L.; Smith, N. A.; and Zhang, C. 2024 · 2024
Closest in time.
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ChatGPT’s One-year Anniversary: Are Open-Source Large Language Models Catching up?
Chen, H.; Jiao, F.; Li, X.; Qin, C.; Ravaut, M.; Zhao, R.; Xiong, C.; and Joty, S. 2024 · 2024
Cited alongside, same era.
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
Fan, C.; Liu, J.; Zhang, Y.; Wong, E.; Wei, D.; and Liu, S. 2024 · 2024
Cited alongside, same era.
Jogging the Memory of Unlearned Model Through Targeted Relearning Attack
Hu, S.; Fu, Y.; Wu, Z. S.; and Smith, V. 2024 · 2024
Cited alongside, same era.
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Huang, Y.; Gupta, S.; Xia, M.; Li, K.; and Chen, D. 2024 · 2024
Cited alongside, same era.
AI Alignment: A Comprehensive Survey
Ji, J.; Qiu, T.; Chen, B.; Zhang, B.; Lou, H.; Wang, K.; Duan, Y.; He, Z.; Zhou, J.; Zhang, Z.; Zeng, F.; Ng, K. Y.; Dai, J.; Pan, X.; O’Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W. 2024 · 2024
Cited alongside, same era.
RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language Models
Jin, Z.; Cao, P.; Wang, C.; He, Z.; Yuan, H.; Li, J.; Chen, Y.; Liu, K.; and Zhao, J. 2024 · 2024
Cited alongside, same era.
Rethinking Machine Unlearning for Large Language Models
Liu, S.; Yao, Y.; Jia, J.; Casper, S.; Baracaldo, N.; Hase, P.; Yao, Y.; Liu, C. Y.; Xu, X.; Li, H.; Varshney, K. R.; Bansal, M.; Koyejo, S.; and Liu, Y. 2024a
Cited in the paper.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2024b
Cited in the paper.
Shumailov, I.; Hayes, J.; Triantafillou, E.; Ortiz-Jimenez, G.; Papernot, N.; Jagielski, M.; Yona, I.; Howard, H.; and Bagdasaryan, E. 2024 · 2024
Closest in time.
Efficient Adversarial Training in LLMs with Continuous Attacks
Xhonneux, S.; Sordoni, A.; Günnemann, S.; Gidel, G.; and Schwinn, L. 2024 · 2024
Closest in time.
Large Language Model Unlearning
Yao, Y.; Xu, X.; and Liu, Y. 2024 · 2024
Closest in time.
Low-Resource Languages Jailbreak GPT-4
Yong, Z.-X.; Menghini, C.; and Bach, S. H. 2024 · 2024
Closest in time.
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Zhang, R.; Lin, L.; Bai, Y.; and Mei, S. 2024 · 2024
Closest in time.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2024
Closest in time.