Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown to pose social and ethical risks such as generating toxic language or facilitating malicious use of hazardous knowledge.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2015
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L · 2019
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Ginart, A., Guan, M., Valiant, G., and Zou, J. Y · 2019
Earlier work this paper cites.
PubmedQA: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Graves, L., Nagisetty, V., and Ganesh, V · 2020
Earlier work this paper cites.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Approximate data deletion from machine learning models
Izzo, Z., Smart, M. A., Chaudhuri, K., and Zou, J · 2021
Earlier work this paper cites.
GeDi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2021
Earlier work this paper cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Unrolling sgd: Understanding factors influencing machine unlearning
Thudi, A., Deza, G., Chandrasekaran, V., and Papernot, N · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2023
Later among the works it cites.
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Kassem, A., Mahmoud, O., and Saad, S · 2023
Later among the works it cites.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C. L., Phang, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Later among the works it cites.
Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools, 2023
Sandbrink, J. B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving alignment of dialogue agents via targeted human judgements, 2022
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Cited alongside, same era.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
QUARK: Controllable text generation with reinforced unlearning
Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y · 2022
Cited alongside, same era.
A survey of machine unlearning, 2022
Nguyen, T. T., Huynh, T. T., Nguyen, P. L., Liew, A. W.-C., Yin, H., and Nguyen, Q. V. H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., and Gabriel, I · 2022
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Later among the works it cites.
KGA: A general machine unlearning framework based on knowledge gap alignment
Wang, L., Chen, T., Yuan, W., Zeng, X., Wong, K.-F., and Yin, H · 2023
Later among the works it cites.
Jailbroken: How does LLM safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Ties-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operation
Zhang, J., Chen, S., Liu, J., and He, J · 2023
Later among the works it cites.
Llm agents can autonomously hack websites, 2024
Fang, R., Bindu, R., Gupta, A., Zhan, Q., and Kang, D · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning, 2024
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., Tamirisa, R., Bharathi, B., Khoja, A., Zhao, Z., Herbert-Voss, A., Breuer, C. B., Marks, S., Patel, O., Zou, A., Mazeika, M., Wang, Z., Oswal, P., Liu, W., Hunt, A. A., Tienken-Harder, J., Shih, K. Y., Talley, K., Guan, J., Kaplan, R., Steneker, I., Campbell, D., Jokubaitis, B., Levinson, A., Wang, J., Qian, W., Karmakar, K. K., Basart, S., Fitz, S., Levine, M., Kumaraguru, P., Tupakula, U., Varadharajan, V., Shoshitaishvili, Y., Ba, J., Esvelt, K. M., Wang, A., and Hendrycks, D · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms, 2024
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2024
Closest in time.
Large language model unlearning, 2024
Yao, Y., Xu, X., and Liu, Y · 2024
Closest in time.
Negative preference optimization: From catastrophic collapse to effective unlearning, 2024
Zhang, R., Lin, L., Bai, Y., and Mei, S · 2024
Closest in time.