Fetching the paper…
Reading the bibliography…
Most interpretability research in NLP focuses on understanding the behavior and features of a fully trained model.
Sarthak Jain and Byron C. Wallace · 1902
Earlier work this paper cites.
The scientific method in the science of machine learning
Jessica Zosa Forde and Michela Paganini · 1904
Earlier work this paper cites.
SGD on Neural Networks Learns Functions of Increasing Complexity
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak · 1905
Earlier work this paper cites.
What Does BERT Look At? An Analysis of BERT’s Attention, June 2019
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 1906
Earlier work this paper cites.
How the mind works
Steven Pinker · 1997
Earlier work this paper cites.
Treebank-3, 1999
Mitchell P. Marcus, Beatrice Santorini, Mary Ann Marcinkiewicz, and Ann Taylor · 1999
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
The Break-Even Point on Optimization Trajectories of Deep Neural Networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2003
Earlier work this paper cites.
Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber · 2004
Earlier work this paper cites.
What shapes feature representations? Exploring datasets, architectures, and training
Katherine L. Hermann and Andrew K. Lampinen · 2006
Earlier work this paper cites.
The Pitfalls of Simplicity Bias in Neural Networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli · 2006
Earlier work this paper cites.
Pretrained Language Model Embryology: The Birth of ALBERT
David C. Chiang, Sung-Feng Huang, and Hung-yi Lee · 2010
Earlier work this paper cites.
Pareto Probing: Trading Off Accuracy for Complexity
Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell · 2010
Earlier work this paper cites.
Gradient Starvation: A Learning Proclivity in Neural Networks
Mohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron Courville, Doina Precup, and Guillaume Lajoie · 2011
Earlier work this paper cites.
Spectrum of human tails: A report of six cases
Biswanath Mukhopadhyay, Ram M Shukla, Madhumita Mukhopadhyay, Kartik C Mandal, Pankaj Haldar, and Abhijit Benare · 2012
Earlier work this paper cites.
Xiaoxia Wu, Ethan Dyer, and Behnam Neyshabur · 2012
Earlier work this paper cites.
Enhanced English Universal Dependencies: An improved representation for natural language understanding tasks
Sebastian Schuster and Christopher D. Manning · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
We’re the Only Animals With Chins, and No One Knows Why, January 2016
Ed Yong · 2016
Earlier work this paper cites.
A Closer Look at Memorization in Deep Networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani · 2017
Earlier work this paper cites.
Universal Dependencies
Joakim Nivre, Daniel Zeman, Filip Ginter, and Francis Tyers · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability, 2017
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Critical learning periods in deep networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2018
Earlier work this paper cites.
The Mythos of Model Interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C. Lipton · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation, 2018
Ari S. Morcos, Maithra Raghu, and Samy Bengio · 2018
Earlier work this paper cites.
Representation compression and generalization in deep neural networks, 2018
Ravid Shwartz-Ziv, Amichai Painsky, and Naftali Tishby · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Earlier work this paper cites.
Intrinsic dimension of data representations in deep neural networks
Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
The inverse scaling prize, 2022a
Ian McKenzie, Alexander Lyzhov, Alicia Parrish, Ameya Prabhu, Aaron Mueller, Najoung Kim, Sam Bowman, and Ethan Perez · 2022
Later among the works it cites.
Characterizing Intrinsic Compositionality in Transformers with Tree Projections, November 2022
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D. Manning · 2022
Later among the works it cites.
A Mechanistic Interpretability Analysis of Grokking, 2022
Neel Nanda and Tom Lieberum · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, T. J. Henighan, Benjamin Mann, Amanda Askell, Yushi Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, John Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom B. Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Christopher Olah · 2022
Later among the works it cites.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding Learning Dynamics Of Language Models with SVCCA
Naomi Saphra and Adam Lopez · 2019
Cited alongside, same era.
Is Attention Interpretable?
Sofia Serrano and Noah A. Smith · 2019
Cited alongside, same era.
The bitter lesson
Richard Sutton · 2019
Cited alongside, same era.
Guillermo Valle-Pérez, Chico Q. Camargo, and Ard A. Louis · 2019
Cited alongside, same era.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema · 2020
Cited alongside, same era.
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
The multiBERTs: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick · 2022
Later among the works it cites.
Information flow in deep neural networks
Ravid Shwartz-Ziv · 2022
Later among the works it cites.
Aarohi Srivastava et al · 2022
Later among the works it cites.
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models, June 2022
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
Machine learning phase transitions: Connections to the fisher information, 2023
Julian Arnold, Niels Lörch, Flemming Holtorf, and Frank Schäfer · 2023
Closest in time.
Reverse engineering self-supervised learning
Ido Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel, and Yann LeCun · 2023
Closest in time.
Broken neural scaling laws
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger · 2023
Closest in time.
ChatGPT Is a Blurry JPEG of the Web
Ted Chiang · 2023
Closest in time.
Are JPEG and LM similar to each other? If so, in what sense, and is this the real question to ask?
Kyunghyun Cho · 2023
Closest in time.
Latent state models of training dynamics, 2023
Michael Y. Hu, Angelica Chen, Naomi Saphra, and Kyunghyun Cho · 2023
Closest in time.
Calibrated chaos: Variance between runs of neural network training is harmless and inevitable, 2023
Keller Jordan · 2023
Closest in time.
Linear connectivity reveals generalization strategies, 2023
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2023
Closest in time.
How does information bottleneck help deep learning?
Kenji Kawaguchi, Zhun Deng, Xu Ji, and Jiaoyang Huang · 2023
Closest in time.
Omnigrok: Grokking beyond algorithmic data, 2023
Ziming Liu, Eric J. Michaud, and Max Tegmark · 2023
Closest in time.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
Grokking of hierarchical structure in vanilla transformers, 2023
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D. Manning · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Measuring and narrowing the compositionality gap in language models, 2023
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis · 2023
Closest in time.
On the special role of class-selective neurons in early training
Omkar Ranadive, Nikhil Thakurdesai, Ari S. Morcos, Matthew L Leavitt, and Stephane Deny · 2023
Closest in time.
Interpretability Creationism
Naomi Saphra · 2023
Closest in time.
Are Emergent Abilities of Large Language Models a Mirage?, April 2023
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Closest in time.
An observation on Generalization, August 2023
Ilya Sutskever · 2023
Closest in time.
Training trajectories of language models across scales, 2023
Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Ves Stoyanov · 2023
Closest in time.