Fetching the paper…
Reading the bibliography…
Despite substantial efforts, neural network interpretability remains an elusive goal, with previous research failing to provide succinct explanations of most single neurons' impact on the network output.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting”
Nitish Srivastava et al · 1958
Earlier work this paper cites.
“Nonlinear signal processing using neural networks: Prediction and system modelling”, 1987
Alan Lapedes and Robert Farber · 1987
Earlier work this paper cites.
“An information theoretic interpretation to deep neural networks”
Shao-Lun Huang, Xiangxiang Xu, Lizhong Zheng and Gregory Wornell · 1988
Earlier work this paper cites.
“Near Shannon limit performance of low density parity check codes”
David MacKay and Radford Neal · 1997
Earlier work this paper cites.
“Deep hierarchies in the primate visual cortex: What can we learn for computer vision?”
Norbert Kruger et al · 2012
Earlier work this paper cites.
“Representation learning: A review and new perspectives”
Yoshua Bengio, Aaron Courville and Pascal Vincent · 2013
Earlier work this paper cites.
“Deep learning and the information bottleneck principle”
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
“Understanding intermediate layers using linear classifier probes”
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Issues in evaluating semantic spaces using word analogies”
Tal Linzen · 2016
Earlier work this paper cites.
Andrei Rusu et al · 2016
Earlier work this paper cites.
“Towards a rigorous science of interpretable machine learning”
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
“Overcoming catastrophic forgetting in neural networks”
James Kirkpatrick et al · 2017
Earlier work this paper cites.
“Learning without forgetting”
Zhizhong Li and Derek Hoiem · 2017
Earlier work this paper cites.
“Opening the black box of deep neural networks via information”
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
“Attention is all you need”
Ashish Vaswani et al · 2017
Cited alongside, same era.
“Mine: mutual information neural estimation”
Mohamed Belghazi et al · 2018
Cited alongside, same era.
“Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks”
Ruth Fong and Andrea Vedaldi · 2018
Cited alongside, same era.
“The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.”
Zachary Lipton · 2018
Cited alongside, same era.
“Zoom in: An introduction to circuits”
Chris Olah et al · 2020
Later among the works it cites.
“Grounding representation similarity through statistical testing”
Frances Ding, Jean-Stanislas Denain and Jacob Steinhardt · 2021
Later among the works it cites.
“A mathematical framework for transformer circuits”
N Elhage et al · 2021
Later among the works it cites.
“Automating auditing: An ambitious concrete technical research proposal”
Evan Hubinger · 2021
Later among the works it cites.
Nelson Elhage et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Insights on representational similarity in neural networks with canonical correlation”
Ari Morcos, Maithra Raghu and Samy Bengio · 2018
Cited alongside, same era.
“Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients”
Andrew Ross and Finale Doshi-Velez · 2018
Cited alongside, same era.
“Adversarial robustness as a prior for learned representations”
Logan Engstrom et al · 2019
Cited alongside, same era.
“Scaleable input gradient regularization for adversarial robustness”
Chris Finlay and Adam Oberman · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Cited alongside, same era.
“High-dimensional geometry of population responses in visual cortex”
Carsen Stringer et al · 2019
Cited alongside, same era.
“Adversarial robustness on in-and out-distribution improves explainability”
Maximilian Augustin, Alexander Meinke and Matthias Hein · 2020
Cited alongside, same era.
Arna Ghosh, Arnab Mondal, Kumar Agrawal and Blake Richards · 2022
Later among the works it cites.
“Planting undetectable backdoors in machine learning models”
Shafi Goldwasser, Michael Kim, Vinod Vaikuntanathan and Or Zamir · 2022
Later among the works it cites.
“Towards Benchmarking Explainable Artificial Intelligence Methods”
Lars Holmberg · 2022
Later among the works it cites.
“Increasing neural network robustness improves match to macaque V1 eigenspectrum, spatial frequency preference and predictivity”
Nathan Kong, Eshed Margalit, Justin Gardner and Anthony Norcia · 2022
Later among the works it cites.
“Capsule networks–a survey”
Mensah Patrick, Adebayo Adekoya, Ayidzoe Mighty and Baagyire Edward · 2022
Later among the works it cites.
“Toward transparent ai: A survey on interpreting the inner structures of deep neural networks”
Tilman R\"aukur, Anson Ho, Stephen Casper and Dylan Hadfield-Menell · 2022
Later among the works it cites.
“Eliciting Latent Predictions from Transformers with the Tuned Lens”
Nora Belrose et al · 2023
Later among the works it cites.
“Language models can explain neurons in language models”,
Steven Bills et al · 2023
Later among the works it cites.
“The scale-invariant covariance spectrum of brain-wide activity”
Zezhen Wang et al · 2023
Later among the works it cites.