Fetching the paper…
Reading the bibliography…
What types of numeric representations emerge in neural systems, and what would a satisfying answer to this question look like? In this work, we interpret Neural Network (NN) solutions to sequence based number tasks using a variety of methods to understand how well we can interpret them through the lens of interpretable Symbolic Algorithms (SAs) -- precise programs describable by rules and typed, mutable variables.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
The Language of Thought
Jerry A. Fodor · 1975
Earlier work this paper cites.
Physical symbol systems
Allen Newell · 1980
Earlier work this paper cites.
Computation and cognition: Issues in the foundations of cognitive science
Zenon W. Pylyshyn · 1980
Earlier work this paper cites.
The knowledge level
Allen Newell · 1982
Earlier work this paper cites.
Parallel Distributed Processing. Volume 2: Psychological and Biological Models
J. L. McClelland, D. E. Rumelhart, and PDP Research Group (eds.) · 1986
Earlier work this paper cites.
Parallel Distributed Processing. Volume 1: Foundations
D. E. Rumelhart, J. L. McClelland, and PDP Research Group (eds.) · 1986
Earlier work this paper cites.
Psychosemantics: The Problem of Meaning in the Philosophy of Mind
Jerry A. Fodor · 1987
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A. Fodor and Zenon W. Pylyshyn · 1988
Earlier work this paper cites.
On the proper treatment of connectionism
Paul Smolensky · 1988
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Numerical cognition without words: Evidence from Amazonia
Peter Gordon · 2004
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
An Introduction to Causal Inference
Judea Pearl · 2010
Earlier work this paper cites.
A Number Sense as an Emergent Property of the Manipulating Brain
Neehar Kondapaneni and Pietro Perona · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Development of numerical cognition in children and artificial systems: a review of the current knowledge and proposals for multi-disciplinary research
Alessandro Di Nuovo and Tim Jay · 2018
Cited alongside, same era.
Can a recurrent neural network learn to count things?
M. Fang, Z. Zhou, S. Chen, and J. L. McClelland · 2018
Cited alongside, same era.
Deep learning: A critical appraisal, 2018
Gary Marcus · 2018
Cited alongside, same era.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Later among the works it cites.
Transformer language models without positional encodings still learn positional information, 2022
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy · 2022
Later among the works it cites.
The neural race reduction: Dynamics of abstraction in gated networks
Andrew M. Saxe, Shagun Sodhani, and Sam Lewallen · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Formal and empirical studies of counting behaviour in relu rnns
Nadine El-Naggar, Andrew Ryzhikov, Laure Daviaud, Pranava Madhyastha, and Tillman Weyde · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interpretable counting for visual question answering
Alexander Trott, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
On the practical computational power of finite precision rnns for language recognition, 2018
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2018
Cited alongside, same era.
Learning to count objects in natural images for visual question answering
Yan Zhang, Jonathon Hare, and Adam Prügel-Bennett · 2018
Cited alongside, same era.
Developing the knowledge of number digits in a child-like robot
Alessandro Di Nuovo and James L. McClelland · 2019
Cited alongside, same era.
Number detectors spontaneously emerge in a deep neural network designed for visual object recognition
Khaled Nasr, Pooja Viswanathan, and Andreas Nieder · 2019
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Later among the works it cites.
Finding alignments between interpretable causal variables and distributed neural representations, 2023
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D. Goodman · 2023
Later among the works it cites.
Dissecting recall of factual associations in auto-regressive language models, 2023
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Later among the works it cites.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Later among the works it cites.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Later among the works it cites.
Distributed representations: Composition & superposition
Chris Olah · 2023
Later among the works it cites.
Polysemanticity and capacity in neural networks, 2023
Adam Scherlis, Kshitij Sachan, Adam S. Jermyn, Joe Benton, and Buck Shlegeris · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Freya Behrens, Luca Biggio, and Lenka Zdeborová · 2024
Later among the works it cites.
The heuristic core: Understanding subnetwork generalization in pretrained language models, 2024
Adithya Bhaskar, Dan Friedman, and Danqi Chen · 2024
Later among the works it cites.
Róbert Csordás, Christopher Potts, Christopher D. Manning, and Atticus Geiger · 2024
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca, 2024
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman · 2024
Later among the works it cites.