Fetching the paper…
Reading the bibliography…
The LLM unlearning technique has recently been introduced to comply with data regulations and address the safety and ethical concerns of LLMs by removing the undesired data-model influence.
Randomized smoothing for stochastic optimization
Duchi, J. C., Bartlett, P. L., and Wainwright, M. J · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Nichol, A · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J., Rosenfeld, E., and Kolter, Z · 2019
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Ginart, A., Guan, M., Valiant, G., and Zou, J. Y · 2019
Earlier work this paper cites.
Visualizing and understanding the effectiveness of bert
Hao, Y., Dong, L., Wei, F., and Xu, K · 2019
Earlier work this paper cites.
Robustness via curvature regularization, and vice versa
Moosavi-Dezfooli, S.-M., Fawzi, A., Uesato, J., and Frossard, P · 2019
Earlier work this paper cites.
Robust overfitting may be mitigated by properly learned smoothening
Chen, T., Zhang, Z., Liu, S., Chang, S., and Wang, Z · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Scaleable input gradient regularization for adversarial robustness
Finlay, C. and Oberman, A. M · 2021
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Earlier work this paper cites.
Machine unlearning via algorithmic stability
Ullah, E., Mai, T., Rao, A., Rossi, R. A., and Arora, R · 2021
Cited alongside, same era.
Towards understanding sharpness-aware minimization
Andriushchenko, M. and Flammarion, N · 2022
Cited alongside, same era.
Sharpness-aware training for free
Du, J., Zhou, D., Feng, J., Tan, V., and Zhou, J. T · 2022
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
Weight perturbation as defense against adversarial word substitutions
Xu, J., Li, L., Zhang, J., Zheng, X., Chang, K.-W., Hsieh, C.-J., and Huang, X.-J · 2022
Towards unbounded machine unlearning
Kurmanji, M., Triantafillou, P., Hayes, J., and Triantafillou, E · 2024
Later among the works it cites.
The WMDP benchmark: Measuring and reducing malicious use with unlearning
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., Tamirisa, R., Bharathi, B., Herbert-Voss, A., Breuer, C. B., Zou, A., Mazeika, M., Wang, Z., Oswal, P., Lin, W., Hunt, A. A., Tienken-Harder, J., Shih, K. Y., Talley, K., Guan, J., Steneker, I., Campbell, D., Jokubaitis, B., Basart, S., Fitz, S., Kumaraguru, P., Karmakar, K. K., Tupakula, U., Varadharajan, V., Shoshitaishvili, Y., Ba, J., Esvelt, K. M., Wang, A., and Hendrycks, D · 2024
Later among the works it cites.
Towards safer large language models through machine unlearning
Liu, Z., Dou, G., Tan, Z., Tian, Y., and Jiang, M · 2024
Later among the works it cites.
An adversarial perspective on machine unlearning for ai safety
Łucki, J., Wei, B., Huang, Y., Henderson, P., Tramèr, F., and Rando, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zan, C., Ding, L., Shen, L., Cao, Y., Liu, W., and Tao, D · 2022
Cited alongside, same era.
Boundary unlearning: Rapid forgetting of deep networks via shifting the decision boundary
Chen, M., Gao, W., Liu, G., Peng, K., and Wang, C · 2023
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms, 2023
Eldan, R. and Russinovich, M · 2023
Cited alongside, same era.
In-context unlearning: Language models as few shot unlearners
Pawelczyk, M., Neel, S., and Lakkaraju, H · 2023
Cited alongside, same era.
Sharpness-aware minimization alone can improve adversarial robustness
Wei, Z., Zhu, J., and Zhang, Y · 2023
Cited alongside, same era.
Depn: Detecting and editing privacy neurons in pretrained language models
Wu, X., Li, J., Xu, M., Dong, W., Wu, S., Bian, C., and Xiong, D · 2023
Cited alongside, same era.
Lynch, A., Guo, P., Ewart, A., Casper, S., and Hadfield-Menell, D · 2024
Later among the works it cites.
TOFU: A task of fictitious unlearning for LLMs
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Patil, V., Hase, P., and Bansal, M · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Latent adversarial training improves robustness to persistent harmful behaviors in llms
Sheshadri, A., Ewart, A., Guo, P., Lynch, A., Wu, C., Hebbar, V., Sleight, H., Stickland, A. C., Perez, E., Hadfield-Menell, D., et al · 2024
Later among the works it cites.
Muse: Machine unlearning six-way evaluation for language models
Shi, W., Lee, J., Huang, Y., Malladi, S., Zhao, J., Holtzman, A., Liu, D., Zettlemoyer, L., Smith, N. A., and Zhang, C · 2024
Later among the works it cites.
Ununlearning: Unlearning is not sufficient for content regulation in advanced generative ai
Shumailov, I., Hayes, J., Triantafillou, E., Ortiz-Jimenez, G., Papernot, N., Jagielski, M., Yona, I., Howard, H., and Bagdasaryan, E · 2024
Later among the works it cites.
Tamper-resistant safeguards for open-weight llms
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms
Thaker, P., Maurya, Y., and Smith, V · 2024
Later among the works it cites.
Flrt: Fluent student-teacher redteaming
Thompson, T. B. and Sklar, M · 2024
Later among the works it cites.
Large language model unlearning
Yao, Y., Xu, X., and Liu, Y · 2024
Later among the works it cites.
When will gradient regularization be harmful?
Zhao, Y., Zhang, H., and Hu, X · 2024
Later among the works it cites.
Open problems in machine unlearning for ai safety
Barez, F., Fu, T., Prabhu, A., Casper, S., Sanyal, A., Bibi, A., O’Gara, A., Kirk, R., Bucknall, B., Fist, T., et al · 2025
Closest in time.
Safety alignment should be made more than just a few tokens deep
Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P · 2025
Closest in time.