Fetching the paper…
Reading the bibliography…
A common method to study deep learning systems is to use simplified model representations--for example, using singular value decomposition to visualize the model's hidden states in a lower dimensional space.
A theory of the learnable
Valiant, L. G · 1984
Earlier work this paper cites.
Perspectives on cognitive neuroscience
Churchland, P. S. and Sejnowski, T. J · 1988
Earlier work this paper cites.
Extraction of rules from discrete-time recurrent neural networks
Omlin, C. W. and Giles, C. L · 1996
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Segregation of object and background motion in the retina
Ölveczky, B. P., Baccus, S. A., and Meister, M · 2003
Earlier work this paper cites.
Rule extraction from recurrent neural networks: A taxonomy and review
Jacobsson, H · 2005
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Implicit regularization in deep learning
Neyshabur, B · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
JAX: Composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Earlier work this paper cites.
Extracting automata from recurrent neural networks using queries and counterexamples
Weiss, G., Goldberg, Y., and Yahav, E · 2018
Earlier work this paper cites.
Non-vacuous generalization bounds at the imagenet scale: A PAC-Bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Belinkov, Y. and Glass, J · 2019
Earlier work this paper cites.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Earlier work this paper cites.
What does BERT look at? An analysis of BERT’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 2019
Earlier work this paper cites.
Interpretation of neural networks is fragile
Ghorbani, A., Abid, A., and Zou, J · 2019
Earlier work this paper cites.
What do compressed deep neural networks forget?
Hooker, S., Courville, A., Clark, G., Dauphin, Y., and Frome, A · 2019
Earlier work this paper cites.
CodeSearchNet challenge: Evaluating the state of semantic code search
Husain, H., Wu, H.-H., Gazit, T., Allamanis, M., and Brockschmidt, M · 2019
Earlier work this paper cites.
Attention is not explanation
Jain, S. and Wallace, B. C · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Visualizing attention in Transformer based language models
Vig, J · 2019
Cited alongside, same era.
Attention is not not explanation
Wiegreffe, S. and Pinter, Y · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
Bau, D., Zhu, J.-Y., Strobelt, H., Lapedriza, A., Zhou, B., and Torralba, A · 2020
Cited alongside, same era.
How can self-attention networks recognize Dyck-n languages?
Ebrahimi, J., Gelda, D., and Zhang, W · 2020
Cited alongside, same era.
How far does BERT look at: Distance-based clustering and analysis of BERT’s attention
Guan, Y., Leng, J., Li, C., Chen, Q., and Guo, M · 2020
Cited alongside, same era.
Haiku: Sonnet for JAX, 2020
Hennigan, T., Cai, T., Norman, T., Martens, L., and Babuschkin, I · 2020
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Closest in time.
A toy model of universality: Reverse engineering how networks learn group operations
Chughtai, B., Chan, L., and Nanda, N · 2023
Closest in time.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Closest in time.
Discovering variable binding circuitry with desiderata
Davies, X., Nadeau, M., Prakash, N., Shaham, T. R., and Bau, D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
RNNs can generate bounded hierarchical languages with optimal memory
Hewitt, J., Hahn, M., Ganguli, S., Liang, P., and Manning, C. D · 2020
Cited alongside, same era.
The EOS decision and length extrapolation
Newman, B., Hewitt, J., Liang, P., and Manning, C. D · 2020
Cited alongside, same era.
Learning to deceive with attention-based explanations
Pruthi, D., Gupta, M., Dhingra, B., Neubig, G., and Lipton, Z. C · 2020
Cited alongside, same era.
Causal mediation analysis for interpreting neural NLP: The case of gender bias
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Sakenis, S., Huang, J., Singer, Y., and Shieber, S · 2020
Cited alongside, same era.
An interpretability illusion for BERT
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Cited alongside, same era.
Dar, Y., Muthukumar, V., and Baraniuk, R. G · 2021
Cited alongside, same era.
Closest in time.
Dissecting recall of factual associations in auto-regressive language models
Geva, M., Bastings, J., Filippova, K., and Globerson, A · 2023
Closest in time.
Localizing model behavior with path patching
Goldowsky-Dill, N., MacLeod, C., Sato, L., and Arora, A · 2023
Closest in time.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Hanna, M., Liu, O., and Variengien, A · 2023
Closest in time.
Does localization inform editing? Surprising differences in causality-based localization vs. knowledge editing in language models
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2023
Closest in time.
The impact of positional encoding on length generalization in Transformers
Kazemnejad, A., Padhi, I., Natesan, K., Das, P., and Reddy, S · 2023
Closest in time.
Lieberum, T., Rahtz, M., Kramár, J., Irving, G., Shah, R., and Mikulik, V · 2023
Closest in time.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
Merrill, W., Tsilivis, N., and Shukla, A · 2023
Closest in time.
Grokking of hierarchical structure in vanilla transformers
Murty, S., Sharma, P., Andreas, J., and Manning, C · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
Do machine learning models memorize or generalize?
Pearce, A., Ghandeharioun, A., Hussein, N., Thain, N., Wattenberg, M., and Dixon, L · 2023
Closest in time.
Understanding arithmetic reasoning in language models using causal mediation analysis
Stolfo, A., Belinkov, Y., and Sachan, M · 2023
Closest in time.
Neurons in large language models: Dead, n-gram, positional
Voita, E., Ferrando, J., and Nalmpantis, C · 2023
Closest in time.
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Closest in time.
Transformers are uninterpretable with myopic methods: A case study with bounded Dyck grammars
Wen, K., Li, Y., Liu, B., and Risteski, A · 2023
Closest in time.
A comprehensive study on post-training quantization for large language models
Yao, Z., Li, C., Wu, X., Youn, S., and He, Y · 2023
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Closest in time.
Learned feature representations are biased by complexity, learning order, position, and more
Lampinen, A. K., Chan, S. C., and Hermann, K · 2024
Closest in time.