Fetching the paper…
Reading the bibliography…
The concept of localization in LLMs is often mentioned in prior work; however, methods for localization have never been systematically and directly evaluated.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I. Levenshtein. 1965 · 1965
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork. 1992 · 1992
Earlier work this paper cites.
Introducing the enron corpus
Bryan Klimt and Yiming Yang. 2004 · 2004
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang. 2015 · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2016 · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017 · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Editable neural networks
Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko. 2020 · 2020
Cited alongside, same era.
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021 · 2021
Cited alongside, same era.
Low-complexity probing via finding subnetworks
Steven Cao, Victor Sanh, and Alexander Rush. 2021 · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
EarlyBERT: Efficient BERT training via early-bird lottery tickets
Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, and Jingjing Liu. 2021 · 2021
Cited alongside, same era.
Robust lottery tickets for pre-trained language models
Rui Zheng, Bao Rong, Yuhao Zhou, Di Liang, Sirui Wang, Wei Wu, Tao Gui, Qi Zhang, and Xuanjing Huang. 2022 · 2022
Later among the works it cites.
Discovering knowledge-critical subnetworks in pretrained language models
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss, and Antoine Bosselut. 2023 · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023 · 2023
Closest in time.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2021 · 2021
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2021 · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Does BERT pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron Wallace. 2021 · 2021
Cited alongside, same era.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Discovering language-neutral sub-networks in multilingual language models
Negar Foroutan, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut, and Karl Aberer. 2022 · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022 · 2022
Cited alongside, same era.
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora. 2023 · 2023
Closest in time.
Editing common sense in transformers
Anshita Gupta, Debanjan Mondal, Akshay Sheshadri, Wenlong Zhao, Xiang Li, Sarah Wiegreffe, and Niket Tandon. 2023 · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. 2023 · 2023
Closest in time.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023 · 2023
Closest in time.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023 · 2023
Closest in time.
Do language models plagiarize?
Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee. 2023 · 2023
Closest in time.
Do question answering modeling improvements hold across benchmarks?
Nelson F. Liu, Tony Lee, Robin Jia, and Percy Liang. 2023 · 2023
Closest in time.
Can neural network memorization be localized?
Pratyush Maini, Michael Curtis Mozer, Hanie Sedghi, Zachary Chase Lipton, J Zico Kolter, and Chiyuan Zhang. 2023 · 2023
Closest in time.
Propagating knowledge updates to lms through distillation
Shankar Padmanabhan, Yasumasa Onoe, Michael Zhang, Greg Durrett, and Eunsol Choi. 2023 · 2023
Closest in time.
Task-specific skill localization in fine-tuned language models
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora. 2023 · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023 · 2023
Closest in time.
DEPN: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023 · 2023
Closest in time.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2023 · 2023
Closest in time.