2023

Who's Harry Potter? Approximate Unlearning in LLMs

Eldan, Ronen, Russinovich, Mark

Understand

Large language models (LLMs) are trained on massive internet corpora that often contain copyrighted content.

  • This poses legal and ethical challenges for the developers and users of these models, as well as the original authors and publishers.
  • In this paper, we propose a novel technique for unlearning a subset of the training data from a LLM, without having to retrain it from scratch.
  • We evaluate our technique on the task of unlearning the Harry Potter books from the Llama2-7b model (a generative language model recently open-sourced by Meta).

Built on

Similar

  • Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering

    Original

    Vikas Yadav, Steven Bethard, and Mihai Surdeanu · 2019

    Cited alongside, same era.

  • Hellaswag: Can a machine really finish your sentence?

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019

    Cited alongside, same era.

  • A framework for few-shot language model evaluation, September 2021

    Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021

    Cited alongside, same era.

  • Machine unlearning survey

    Yiwen Jiang, Shenglong Liu, Tao Zhao, Wei Li, and Xianzhou Gao · 2022

    Cited alongside, same era.

  • Knowledge unlearning for mitigating privacy risks in language models

    Original

    Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2022

    Cited alongside, same era.

Then

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…