Fetching the paper…
Reading the bibliography…
Neural network models have achieved high performance on a wide variety of complex tasks, but the algorithms that they implement are notoriously difficult to interpret.
Aspects of the Theory of Syntax
Noam Chomsky · 1965
Earlier work this paper cites.
Learning a nonlinear embedding by preserving class neighbourhood structure
Ruslan Salakhutdinov and Geoff Hinton · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen · 2018
Earlier work this paper cites.
Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
R Thomas McCoy, Robert Frank, and Tal Linzen · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2018
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al · 2018
Earlier work this paper cites.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman · 2018
Earlier work this paper cites.
Analyzing and improving representations with the soft nearest neighbor loss
Nicholas Frosst, Nicolas Papernot, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Yoav Goldberg · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Deepmerge: Classifying high-redshift merging galaxies with deep neural networks
Aleksandra Ćiprijanović, Gregory F Snyder, Brian Nord, and Joshua EG Peek · 2020
Earlier work this paper cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Cited alongside, same era.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Cited alongside, same era.
Winning the lottery with continuous sparsification
Pedro Savarese, Hugo Silva, and Michael Maire · 2020
Cited alongside, same era.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Later among the works it cites.
Is a modular architecture enough?
Sarthak Mittal, Yoshua Bengio, and Guillaume Lajoie · 2022
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Christopher Olah · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, T. J. Henighan, Benjamin Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information-theoretic probing with minimum description length
E Voita and I Titov · 2020
Cited alongside, same era.
Assessing phrasal representation and composition in transformers
Lang Yu and Allyson Ettinger · 2020
Cited alongside, same era.
Low-complexity probing via finding subnetworks
Steven Cao, Victor Sanh, and Alexander M Rush · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Cited alongside, same era.
Later among the works it cites.
Semantic structure in deep learning
Ellie Pavlick · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Leace: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al · 2023
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman · 2023
Closest in time.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model, 2023
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Closest in time.
Language models implement simple word2vec-style vector arithmetic, 2023
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Closest in time.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Closest in time.
Task-specific skill localization in fine-tuned language models
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora · 2023
Closest in time.
Symbols and grounding in large language models
Ellie Pavlick · 2023
Closest in time.
Gold doesn’t always glitter: Spectral removal of linear and nonlinear guarded attribute information
Shun Shao, Yftah Ziser, and Shay B Cohen · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D Goodman · 2023
Closest in time.