Fetching the paper…
Reading the bibliography…
Understanding neural networks is challenging in part because of the dense, continuous nature of their hidden states.
Selected studies of the principle of relative frequency in language
George Kingsley Zipf · 1932
Earlier work this paper cites.
5 The Spandrels of San Marco and the Panglossian Paradigm: A Critique of the Adaptationist Programme
Stephen Jay Gould and Richard C Lewontin · 1979
Earlier work this paper cites.
On the role of scientific thought
Edsger W Dijkstra and Edsger W Dijkstra · 1982
Earlier work this paper cites.
Vector quantization
Robert Gray · 1984
Earlier work this paper cites.
A general framework for parallel distributed processing
David E Rumelhart, Geoffrey E Hinton, James L McClelland, et al · 1986
Earlier work this paper cites.
Sparse distributed memory
Pentti Kanerva · 1988
Earlier work this paper cites.
Parallel distributed processing
David E Rumelhart, James L McClelland, PDP Research Group, et al · 1988
Earlier work this paper cites.
Learning sequential structure in simple recurrent networks
David Servan-Schreiber, Axel Cleeremans, and James McClelland · 1988
Earlier work this paper cites.
Local vs. distributed coding
Simon Thorpe · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Divergence measures based on the shannon entropy
Jianhua Lin · 1991
Earlier work this paper cites.
The exaptive excellence of spandrels as a term and prototype
Stephen Jay Gould · 1997
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by V1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Causation, prediction, and search
Peter Spirtes, Clark N Glymour, and Richard Scheines · 2000
Earlier work this paper cites.
Rule extraction from recurrent neural networks: Ataxonomy and review
Henrik Jacobsson · 2005
Earlier work this paper cites.
Stable signal recovery from incomplete and inaccurate measurements
Emmanuel J Candes, Justin K Romberg, and Terence Tao · 2006
Earlier work this paper cites.
Compressed sensing
David L Donoho · 2006
Earlier work this paper cites.
Image denoising via sparse and redundant representations over learned dictionaries
Michael Elad and Michal Aharon · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng · 2006
Earlier work this paper cites.
Sparse coding via thresholding and local competition in neural circuits
Christopher J Rozell, Don H Johnson, Richard G Baraniuk, and Bruno A Olshausen · 2008
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
Jun Zhu and Eric P Xing · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Pedestrian detection with unsupervised multi-stage feature learning
Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala, and Yann LeCun · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Composite quantization for approximate nearest neighbor search
Ting Zhang, Chao Du, and Jingdong Wang · 2014
Earlier work this paper cites.
Winner-take-all autoencoders
Alireza Makhzani and Brendan J Frey · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Pointer Sentinel Mixture Models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al · 2017
Cited alongside, same era.
Neural discrete representation learning
Transformer Feed-Forward Layers Are Key-Value Memories, 2021
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Later among the works it cites.
Multimodal Neurons in Artificial Neural Networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Later among the works it cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning · 2021
Later among the works it cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2021
Later among the works it cites.
Editing a classifier by rewriting its prediction rules
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aaron van den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Cited alongside, same era.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Ruth Fong and Andrea Vedaldi · 2018
Cited alongside, same era.
Mario Giulianelli, Jacqueline Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Cited alongside, same era.
The Building Blocks of Interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Cited alongside, same era.
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie · 2018
Cited alongside, same era.
Eric Wong, Shibani Santurkar, and Aleksander Madry · 2021
Later among the works it cites.
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu · 2021
Later among the works it cites.
Zeyu Yun, Yubei Chen, Bruno A Olshausen, and Yann LeCun · 2021
Later among the works it cites.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Sid Black, Stella Rose Biderman, Eric Hallahan, Quentin G. Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Martin Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Benqi Wang, and Samuel Weinbach · 2022
Later among the works it cites.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Lawrence Chan, Adrià Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishinskaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas · 2022
Later among the works it cites.
Softmax Linear Units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Later among the works it cites.
Post-hoc Interpretability for Neural NLP: A Survey
Andreas Madsen, Siva Reddy, and Sarath Chandar · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge · 2022
Later among the works it cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Post-hoc concept bottleneck models
Mert Yuksekgonul, Maggie Wang, and James Zou · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Closest in time.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Closest in time.
TinyStories: How Small Can Language Models Be and Still Speak Coherent English?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
Dan Friedman, Alexander Wettig, and Danqi Chen · 2023
Closest in time.
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman · 2023
Closest in time.
Measuring and manipulating knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas · 2023
Closest in time.
Backpack Language Models, 2023
John Hewitt, John Thickstun, Christopher D. Manning, and Percy Liang · 2023
Closest in time.
Seeing is Believing: Brain-Inspired Modular Training for Mechanistic Interpretability, 2023
Ziming Liu, Eric Gan, and Max Tegmark · 2023
Closest in time.
Language Models Implement Simple Word2Vec-style Vector Arithmetic
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick · 2023
Closest in time.
Activation Addition: Steering Language Models Without Optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Closest in time.