Fetching the paper…
Reading the bibliography…
How does the internal computation of a machine learning model transform inputs into predictions? In this paper, we introduce a task called component modeling that aims to address this question.
“Computational graphs and rounding error”
Friedrich Bauer · 1974
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”
Alex Krizhevsky · 2009
Earlier work this paper cites.
“The caltech-ucsd birds-200-2011 dataset”
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona and Serge Belongie · 2011
Earlier work this paper cites.
“Poisoning attacks against support vector machines”
Battista Biggio, Blaine Nelson and Pavel Laskov · 2012
Earlier work this paper cites.
“Deep inside convolutional networks: Visualising image classification models and saliency maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2013
Earlier work this paper cites.
“Visualizing and understanding convolutional networks”
Matthew Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian Goodfellow, Jonathon Shlens and Christian Szegedy · 2015
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Earlier work this paper cites.
“Deep Learning Face Attributes in the Wild”
Ziwei Liu, Ping Luo, Xiaogang Wang and Xiaoou Tang · 2015
Earlier work this paper cites.
“Understanding intermediate layers using linear classifier probes”
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
“" Why should I trust you?" Explaining the predictions of any classifier”
Marco Ribeiro, Sameer Singh and Carlos Guestrin · 2016
Earlier work this paper cites.
“Network dissection: Quantifying interpretability of deep visual representations”
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva and Antonio Torralba · 2017
Earlier work this paper cites.
“Badnets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain”
Tianyu Gu, Brendan Dolan-Gavitt and Siddharth Garg · 2017
Earlier work this paper cites.
“Learning to generate reviews and discovering sentiment”
Alec Radford, Rafal Jozefowicz and Ilya Sutskever · 2017
Earlier work this paper cites.
“Axiomatic attribution for deep networks”
Mukund Sundararajan, Ankur Taly and Qiqi Yan · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Places: A 10 million image database for scene recognition”
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva and Antonio Torralba · 2017
Earlier work this paper cites.
“Gender shades: Intersectional accuracy disparities in commercial gender classification”
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
“Deep RNNs encode soft hierarchical syntax”
Terra Blevins, Omer Levy and Luke Zettlemoyer · 2018
Earlier work this paper cites.
“Recognition in terra incognita”
Sara Beery, Grant Van and Pietro Perona · 2018
Earlier work this paper cites.
Kedar Dhamdhere, Mukund Sundararajan and Qiqi Yan · 2018
Earlier work this paper cites.
“A benchmark for interpretability methods in deep neural networks”
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans and Been Kim · 2018
Earlier work this paper cites.
“Influence-directed explanations for deep convolutional networks”
Klas Leino, Shayak Sen, Anupam Datta, Matt Fredrikson and Linyi Li · 2018
Earlier work this paper cites.
“The Building Blocks of Interpretability”
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye and Alexander Mordvintsev · 2018
Earlier work this paper cites.
“Revisiting the importance of individual units in cnns via ablation”
Bolei Zhou, Yiyou Sun, David Bau and Antonio Torralba · 2018
Earlier work this paper cites.
“BoolQ: Exploring the surprising difficulty of natural yes/no questions”
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins and Kristina Toutanova · 2019
Earlier work this paper cites.
“What is one grain of sand in the desert? analyzing individual neurons in deep nlp models”
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau and James Glass · 2019
Earlier work this paper cites.
“ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.”
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix. Wichmann and Wieland Brendel · 2019
Earlier work this paper cites.
“Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations”
Dan Hendrycks and Thomas. Dietterich · 2019
Earlier work this paper cites.
“Designing and interpreting probes with control tasks”
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
“MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs”
Alistair Johnson, Tom Pollard, Nathaniel Greenbaum, Matthew Lungren, Chih-ying Deng, Yifan Peng, Zhiyong Lu, Roger Mark, Seth Berkowitz and Steven Horng · 2019
Earlier work this paper cites.
“Sgd on neural networks learns functions of increasing complexity”
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak and Haofeng Zhang · 2019
Earlier work this paper cites.
“Similarity of Neural Network Representations Revisited”
Simon Kornblith, Mohammad Norouzi, Honglak Lee and Geoffrey Hinton · 2019
Earlier work this paper cites.
“The emergence of number and syntax units in LSTM language models”
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene and Marco Baroni · 2019
Earlier work this paper cites.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“The woman worked as a babysitter: On biases in language generation”
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan and Nanyun Peng · 2019
Earlier work this paper cites.
“Learning robust global representations by penalizing local predictive power”
Haohan Wang, Songwei Ge, Eric Xing and Zachary Lipton · 2019
Earlier work this paper cites.
“Rewriting a deep generative model”
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu and Antonio Torralba · 2020
Earlier work this paper cites.
“Understanding the role of individual units in a deep neural network”
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou and Antonio Torralba · 2020
Earlier work this paper cites.
“Curve detectors”
Nick Cammarata, Gabriel Goh, Shan Carter, Ludwig Schubert, Michael Petrov and Chris Olah · 2020
Earlier work this paper cites.
“Analyzing individual neurons in pre-trained language models”
Nadir Durrani, Hassan Sajjad, Fahim Dalvi and Yonatan Belinkov · 2020
Cited alongside, same era.
“What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation”
Vitaly Feldman and Chiyuan Zhang · 2020
Cited alongside, same era.
“Model patching: Closing the subgroup performance gap with data augmentation”
Karan Goel, Albert Gu, Yixuan Li and Christopher Ré · 2020
Cited alongside, same era.
“Shortcut learning in deep neural networks”
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge and Felix Wichmann · 2020
Cited alongside, same era.
“Neuron shapley: Discovering the responsible neurons”
Amirata Ghorbani and James Zou · 2020
Cited alongside, same era.
“In-context learning and induction heads”
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai and Anna Chen · 2022
Later among the works it cites.
“Clip-dissect: Automatic description of neuron representations in deep vision networks”
Tuomas Oikarinen and Tsui-Wei Weng · 2022
Later among the works it cites.
“Linear adversarial concept erasure”
Shauli Ravfogel, Michael Twiton, Yoav Goldberg and Ryan Cotterell · 2022
Later among the works it cites.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small”
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris and Jacob Steinhardt · 2022
Later among the works it cites.
“Language models can explain neurons in language models”
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu and William Saunders · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya and Siqi Yan · 2020
Cited alongside, same era.
“Celeb-df: A large-scale challenging dataset for deepfake forensics”
Yuezun Li, Xin Yang, Pu Sun, Honggang Qi and Siwei Lyu · 2020
Cited alongside, same era.
“Compositional explanations of neurons”
Jesse Mu and Jacob Andreas · 2020
Cited alongside, same era.
“An Overview of Early Vision in InceptionV1”
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov and Shan Carter · 2020
Cited alongside, same era.
“Zoom In: An Introduction to Circuits”
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov and Shan Carter · 2020
Cited alongside, same era.
“Hidden stratification causes clinically meaningful failures in machine learning for medical imaging”
Lauren Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro and Christopher Ré · 2020
Cited alongside, same era.
“Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization”
Shiori Sagawa, Pang Koh, Tatsunori. Hashimoto and Percy Liang · 2020
Cited alongside, same era.
Later among the works it cites.
“Discovering Knowledge-Critical Subnetworks in Pretrained Language Models”
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss and Antoine Bosselut · 2023
Later among the works it cites.
“Robustness of edited neural networks”
Davis Brown, Charles Godfrey, Cody. Nizinski, Jonathan Tu and Henry Kvinge · 2023
Later among the works it cites.
“Evaluating the Ripple Effects of Knowledge Editing in Language Models”
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson and Mor Geva · 2023
Later among the works it cites.
“Towards automated circuit discovery for mechanistic interpretability”
Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim and Adrià Garriga-Alonso · 2023
Later among the works it cites.
“Do Localization Methods Actually Localize Memorized Data in LLMs?”
Ting-Yun Chang, Jesse Thomason and Robin Jia · 2023
Later among the works it cites.
“Interpreting and Controlling Vision Foundation Models via Text Explanations”
Haozhe Chen, Junfeng Yang, Carl Vondrick and Chengzhi Mao · 2023
Later among the works it cites.
“TinyStories: How Small Can Language Models Be and Still Speak Coherent English?”
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
“Interpreting CLIP’s Image Representation via Text-Based Decomposition”
Yossi Gandelsman, Alexei Efros and Jacob Steinhardt · 2023
Later among the works it cites.
“Erasing concepts from diffusion models”
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman and David Bau · 2023
Later among the works it cites.
“Localizing model behavior with path patching”
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato and Aryaman Arora · 2023
Later among the works it cites.
“Causal abstraction for faithful model interpretation”
Atticus Geiger, Chris Potts and Thomas Icard · 2023
Later among the works it cites.
“A framework for few-shot language model evaluation”
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang and Andy Zou · 2023
Later among the works it cites.
“The journey, not the destination: How data guides diffusion models”
Kristian Georgiev, Joshua Vendrow, Hadi Salman, Sung Park and Aleksander Madry · 2023
Later among the works it cites.
“Don’t trust your eyes: on the (un) reliability of feature visualizations”
Robert Geirhos, Roland Zimmermann, Blair Bilodeau, Wieland Brendel and Been Kim · 2023
Later among the works it cites.
Peter Hase, Mohit Bansal, Been Kim and Asma Ghandeharioun · 2023
Later among the works it cites.
“Rigorously Assessing Natural Language Explanations of Neurons”
Jing Huang, Atticus Geiger, Karel D’Oosterlinck, Zhengxuan Wu and Christopher Potts · 2023
Later among the works it cites.
“On the Foundations of Shortcut Learning”
Katherine Hermann, Hossein Mobahi, Thomas Fel and Michael Mozer · 2023
Later among the works it cites.
“Transformer-Patcher: One Mistake worth One Neuron”
Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong and Zhang Xiong · 2023
Later among the works it cites.
“Phi-2: The surprising power of small language models”
Mojan Javaheripi and Sébastien Bubeck · 2023
Later among the works it cites.
“Textbooks are all you need ii: phi-1.5 technical report”
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del, Suriya Gunasekar and Yin Lee · 2023
Later among the works it cites.
“Can Neural Network Memorization Be Localized?”
Pratyush Maini, Michael Mozer, Hanie Sedghi, Zachary Lipton, J Kolter and Chiyuan Zhang · 2023
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith and Jacob Steinhardt · 2023
Later among the works it cites.
“TRAK: Attributing Model Behavior at Scale”
Sung Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc and Aleksander Madry · 2023
Later among the works it cites.
“Task-Specific Skill Localization in Fine-tuned Language Models”
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao and Sanjeev Arora · 2023
Later among the works it cites.
“Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization”
Elan Rosenfeld and Andrej Risteski · 2023
Later among the works it cites.
“Understanding Arithmetic Reasoning in Language Models using Causal Mediation Analysis”
Alessandro Stolfo, Yonatan Belinkov and Mrinmaya Sachan · 2023
Later among the works it cites.
“Modeldiff: A framework for comparing learning algorithms”
Harshay Shah, Sung Park, Andrew Ilyas and Aleksander Madry · 2023
Later among the works it cites.
“Attribution Patching Outperforms Automated Circuit Discovery”
Aaquib Syed, Can Rager and Arthur Conmy · 2023
Later among the works it cites.
“Dataset interfaces: Diagnosing model failures using controllable counterfactual generation”
Joshua Vendrow, Saachi Jain, Logan Engstrom and Aleksander Madry · 2023
Later among the works it cites.
“Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars”
Kaiyue Wen, Yuchen Li, Bingbin Liu and Andrej Risteski · 2023
Later among the works it cites.
“A comprehensive survey of forgetting in deep learning beyond continual learning”
Zhenyi Wang, Enneng Yang, Li Shen and Heng Huang · 2023
Later among the works it cites.
“Towards Best Practices of Activation Patching in Language Models: Metrics and Methods”
Fred Zhang and Neel Nanda · 2023
Later among the works it cites.
“Representation engineering: A top-down approach to ai transparency”
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika and Ann-Kathrin Dombrowski · 2023
Later among the works it cites.
“Intriguing Properties of Data Attribution on Diffusion Models”
Xiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang and Min Lin · 2023
Later among the works it cites.
“AtP*: An efficient and scalable method for localizing LLM behaviour to components”
János Kramár, Tom Lieberum, Rohin Shah and Neel Nanda · 2024
Closest in time.