Fetching the paper…
Reading the bibliography…
Through considerable effort and intuition, several recent works have reverse-engineered nontrivial behaviors of transformer models.
“Optimal brain damage”, 1989
Yann LeCun, John Denker and Sara Solla · 1989
Earlier work this paper cites.
“Second order derivatives for network pruning: Optimal brain surgeon”, 1992
Babak Hassibi and David Stork · 1992
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2001
Earlier work this paper cites.
“Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias”
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer and Stuart Shieber · 2004
Earlier work this paper cites.
“An introduction to ROC analysis” ROC Analysis in Pattern Recognition
Tom Fawcett · 2005
Earlier work this paper cites.
“Compositional Explanations of Neurons”, 2021
Jesse Mu and Jacob Andreas · 2006
Earlier work this paper cites.
“Rewriting a Deep Generative Model”
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu and Antonio Torralba · 2007
Earlier work this paper cites.
“Causality”
Judea Pearl · 2009
Earlier work this paper cites.
“An overview of 11 proposals for building safe advanced AI”, 2020
Evan Hubinger · 2012
Earlier work this paper cites.
“The Mythos of Model Interpretability”
Zachary. Lipton · 2016
Earlier work this paper cites.
“Residual Networks Behave Like Ensembles of Relatively Shallow Networks”
Andreas Veit, Michael. Wilber and Serge. Belongie · 2016
Earlier work this paper cites.
“Interpretable Explanations of Black Boxes by Meaningful Perturbation”
Ruth. Fong and Andrea Vedaldi · 2017
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax”, 2017
Eric Jang, Shixiang Gu and Ben Poole · 2017
Earlier work this paper cites.
“Feature Visualization” https://distill.pub/2017/feature-visualization
Chris Olah, Alexander Mordvintsev and Ludwig Schubert · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“The malicious use of artificial intelligence: Forecasting, prevention, and mitigation”
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff and Bobby Filar · 2018
Earlier work this paper cites.
“Learning Sparse Neural Networks through L_0 Regularization”
Christos Louizos, Max Welling and Diederik. Kingma · 2018
Earlier work this paper cites.
“Language Modeling Teaches You More than Translation Does: Lessons Learned Through Auxiliary Syntactic Task Analysis”
Kelly Zhang and Samuel Bowman · 2018
Earlier work this paper cites.
“Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshop”
Afra Alishahi, Grzegorz Chrupała and Tal Linzen · 2019
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“OpenWebText Corpus”, 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick and Stefanie Tellex · 2019
Earlier work this paper cites.
“Are Sixteen Heads Really Better than One?”
Paul Michel, Omer Levy and Graham Neubig · 2019
Earlier work this paper cites.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?”
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
“What’s Hidden in a Randomly Weighted Neural Network?”
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi and Mohammad Rastegari · 2020
Cited alongside, same era.
“Movement Pruning: Adaptive Sparsity by Fine-Tuning”
Victor Sanh, Thomas Wolf and Alexander. Rush · 2020
Cited alongside, same era.
“Structured Pruning of Large Language Models”
Ziheng Wang, Jeremy Wohlwend and Tao Lei · 2020
Cited alongside, same era.
“Analysis of Explainers of Black Box Deep Neural Networks for Computer Vision: A Survey”, 2021, pp. 966–989
Vanessa Buhrmester, David Münch and Michael Arens · 2021
Cited alongside, same era.
“Curve Circuits” https://distill.pub/2020/circuits/curve-circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert and Chris Olah · 2021
Cited alongside, same era.
“Low-Complexity Probing via Finding Subnetworks”
Steven Cao, Victor Sanh and Alexander Rush · 2021
“Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks”
Tilman Räuker, Anson Ho, Stephen Casper and Dylan Hadfield-Menell · 2022
Later among the works it cites.
“Emergent Abilities of Large Language Models”, 2022
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus · 2022
Later among the works it cites.
“Causal Distillation for Language Models”
Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts and Noah Goodman · 2022
Later among the works it cites.
“Language models can explain neurons in language models”, https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html , 2023
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu and William Saunders · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“A Mathematical Framework for Transformer Circuits”
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish and Chris Olah · 2021
Cited alongside, same era.
“CausaLM: Causal Model Explanation Through Counterfactual Language Models”
Amir Feder, Nadav Oved, Uri Shalit and Roi Reichart · 2021
Cited alongside, same era.
“Causal Abstractions of Neural Networks”
Atticus Geiger, Hanson Lu, Thomas Icard and Christopher Potts · 2021
Cited alongside, same era.
“Transformer Feed-Forward Layers Are Key-Value Memories”
Mor Geva, Roei Schuster, Jonathan Berant and Omer Levy · 2021
Cited alongside, same era.
“A survey on neural network interpretability”
Yu Zhang, Peter Tiňo, Aleš Leonardis and Ke Tang · 2021
Cited alongside, same era.
“Causal scrubbing: A method for rigorously testing interpretability hypotheses”, Alignment Forum, 2022
Lawrence Chan, Adria Garriga-Alonso, Nix Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris and Nate Thomas · 2022
Cited alongside, same era.
Bilal Chughtai, Lawrence Chan and Neel Nanda · 2023
Closest in time.
“SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot”
Elias Frantar and Dan Alistarh · 2023
Closest in time.
“Hungry Hungry Hippos: Towards Language Modeling with State Space Models”
Daniel Fu, Tri Dao, Khaled Saab, Armin Thomas, Atri Rudra and Christopher Re · 2023
Closest in time.
“Localizing Model Behavior with Path Patching”, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato and Aryaman Arora · 2023
Closest in time.
“Finding Neurons in a Haystack: Case Studies with Sparse Probing”, 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii and Dimitris Bertsimas · 2023
Closest in time.
Michael Hanna, Ollie Liu and Alexandre Variengien · 2023
Closest in time.
“A circuit for Python docstrings in a 4-layer attention-only transformer”, 2023
Stefan Heimersheim and Jett Janiak · 2023
Closest in time.
“An Overview of Catastrophic AI Risks”, 2023
Dan Hendrycks, Mantas Mazeika and Thomas Woodside · 2023
Closest in time.
“Tracr: Compiled Transformers as a Laboratory for Interpretability”, 2023
David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath and Vladimir Mikulik · 2023
Closest in time.
“Identifying a Preliminary Circuit for Predicting Gendered Pronouns in GPT-2 Small”, 2023
Chris Mathwin, Guillaume Corlouer, Esben Kran, Fazl Barez and Neel Nanda · 2023
Closest in time.
“Copy Suppression: Comprehensively Understanding an Attention Head”, 2023
Callum McDougall, Arthur Conmy, Cody Rushing, Thomas McGrath and Neel Nanda · 2023
Closest in time.
“Attribution Patching: Activation Patching At Industrial Scale”, 2023
Neel Nanda · 2023
Closest in time.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith and Jacob Steinhardt · 2023
Closest in time.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Closest in time.
“Attribution Patching Outperforms Automated Circuit Discovery”, 2023
Aaquib Syed, Can Rager and Arthur Conmy · 2023
Closest in time.
“Linear Representations of Sentiment in Large Language Models”, 2023
Curt Tigges, Oskar Hollinsworth, Atticus Geiger and Neel Nanda · 2023
Closest in time.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris and Jacob Steinhardt · 2023
Closest in time.
“Interpretability at Scale: Identifying Causal Mechanisms in Alpaca”, 2023
Zhengxuan Wu, Atticus Geiger, Christopher Potts and Noah. Goodman · 2023
Closest in time.
“A Survey on Model Compression for Large Language Models”, 2023
Xunyu Zhu, Jian Li, Yong Liu, Can Ma and Weiping Wang · 2023
Closest in time.