Fetching the paper…
Reading the bibliography…
Language models often exhibit behaviors that improve performance on a pre-training objective but harm performance on downstream tasks.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D · 1992
Earlier work this paper cites.
Causality and model abstraction
Iwasaki, Y. and Simon, H. A · 1994
Earlier work this paper cites.
A value for n-person games
Shapley, L. S · 1997
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Object detectors emerge in deep scene cnns, 2015
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2015
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction
Lakkaraju, H., Bach, S. H., and Leskovec, J · 2016
Earlier work this paper cites.
Real time image saliency for black box classifiers
Dabkowski, P. and Gal, Y · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C. and Vedaldi, A · 2017
Earlier work this paper cites.
Learning sparse neural networks through l _ 0 l\_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S.-I · 2017
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation, 2018
Chen, J., Song, L., Wainwright, M. J., and Jordan, M. I · 2018
Earlier work this paper cites.
Manipulating and measuring model interpretability
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Vaughan, J. W., and Wallach, H. M · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Structured pruning of large language models
Wang, Z., Wohlwend, J., and Lei, T · 2019
Earlier work this paper cites.
INVASE: Instance-wise variable selection using neural networks
Yoon, J., Jordon, J., and van der Schaar, M · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Thread: circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Cited alongside, same era.
Understanding global feature contributions with additive importance measures, 2020
Covert, I., Lundberg, S., and Lee, S.-I · 2020
Cited alongside, same era.
Neuron shapley: Discovering the responsible neurons, 2020
Ghorbani, A. and Zou, J · 2020
Cited alongside, same era.
Few-shot backdoor defense using shapley estimation
Guan, J., Tu, Z., He, R., and Tao, D · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Olah, C · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Räukur, T., Ho, A., Casper, S., and Hadfield-Menell, D · 2022
Later among the works it cites.
Learning to summarize from human feedback, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Golatkar, A., Achille, A., and Soatto, S · 2020
Cited alongside, same era.
Detoxify
Hanu, L. and Unitary team · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
Raiders of the lost kek: 3.5 years of augmented 4chan posts from the politically incorrect board, 2020
Papasavva, A., Zannettou, S., Cristofaro, E. D., Stringhini, G., and Blackburn, J · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Cited alongside, same era.
Adversarial training for high-stakes reliability, 2022
Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt-Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N · 2020
Cited alongside, same era.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Cited alongside, same era.
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Later among the works it cites.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Closest in time.
Erasing concepts from diffusion models
Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D · 2023
Closest in time.
Localizing model behavior with path patching
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Closest in time.
Perspective API, 2023
Google · 2023
Closest in time.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Closest in time.
Measuring and manipulating knowledge representations in language models
Hernandez, E., Li, B. Z., and Andreas, J · 2023
Closest in time.
Editing Models with Task Arithmetic, March 2023
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Closest in time.
Attribution patching: Activation patching at industrial scale, 2023
Nanda, N · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Liberum, T., Smith, J., and Steinhardt, J · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.