Fetching the paper…
Reading the bibliography…
The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 1912
Earlier work this paper cites.
Group Equivariant Convolutional Networks
Cohen, T. and Welling, M · 1938
Earlier work this paper cites.
On an elementary proof of some asymptotic formulas in the theory of partitions
Erdos, P · 1942
Earlier work this paper cites.
Outer automorphisms of S 6 S_{6}
Janusz, G. and Rotman, J · 1982
Earlier work this paper cites.
Principal component analysis
Wold, S., Esbensen, K., and Geladi, P · 1987
Earlier work this paper cites.
Group Representations in Probability and Statistics , volume 11 of Institute of Mathematical Statistics Lecture Notes
Diaconis, P · 1988
Earlier work this paper cites.
Representation Theory
Fulton, W. and Harris, Joe · 1991
Earlier work this paper cites.
Fast Fourier Transforms for Symmetric Groups: Theory and Implementation
Clausen, M. and Baum, U · 1993
Earlier work this paper cites.
Enumerating finite groups of given order
Pyber, L · 1993
Earlier work this paper cites.
Abstract Algebra
Dummit, D. S. and Foote, R. M · 2003
Earlier work this paper cites.
Fourier analysis: an introduction
Elias M. Stein, R. S · 2003
Earlier work this paper cites.
Group theoretical methods in machine learning
Kondor, R · 2008
Earlier work this paper cites.
Fourier Theoretic Probabilistic Inference over Permutations
Huang, J., Guestrin, C., and Guibas, L · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps, 2014
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
SnFFT: A Julia Toolkit for Fourier Analysis of Functions over Permutations
Plumb, G., Pachauri, D., Kondor, R., and Singh, V · 2015
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Intersecting Families of Permutations, July 2017
Ellis, D., Friedgut, E., and Pilpel, H · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions, 2017
Lundberg, S. and Lee, S.-I · 2017
Earlier work this paper cites.
Not just a black box: Learning important features through propagating activation differences, 2017
Shrikumar, A., Greenside, P., Shcherbina, A., and Kundaje, A · 2017
Earlier work this paper cites.
On the Generalization of Equivariance and Convolution in Neural Networks to the Action of Compact Groups
Kondor, R. and Trivedi, S · 2018
Cited alongside, same era.
Attention is not explanation
Jain, S. and Wallace, B. C · 2019
Cited alongside, same era.
Sanity checks for saliency maps, 2020
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2020
Cited alongside, same era.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Cited alongside, same era.
Zoom In: An Introduction to Circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Cited alongside, same era.
An interpretability illusion for bert, 2021
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models, 2023
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2023
Closest in time.
Neural Discovery of Permutation Subgroups
Karjol, P., Kashyap, R., and Ap, P · 2023
Closest in time.
Grokking as the Transition from Lazy to Rich Training Dynamics, October 2023
Kumar, T., Bordelon, B., Gershman, S. J., and Pehlevan, C · 2023
Closest in time.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla, 2023
Lieberum, T., Rahtz, M., Kramár, J., Nanda, N., Irving, G., Shah, R., and Mikulik, V · 2023
Closest in time.
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching, 2023
Makelov, A., Lange, G., and Nanda, N · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories, 2021
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Cited alongside, same era.
Transformerlens
Nanda, N. and Bloom, J · 2022
Cited alongside, same era.
In-context Learning and Induction Heads, September 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, January 2022
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Cited alongside, same era.
Emergent abilities of large language models, 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Cited alongside, same era.
Red teaming deep neural networks with feature synthesis tools, 2023
Casper, S., Li, Y., Li, J., Bu, T., Zhang, K., Hariharan, K., and Hadfield-Menell, D · 2023
Cited alongside, same era.
The hydra effect: Emergent self-repair in language model computations, 2023
McGrath, T., Rahtz, M., Kramar, J., Mikulik, V., and Legg, S · 2023
Closest in time.
Locating and editing factual associations in gpt, 2023
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2023
Closest in time.
A Tale of Two Circuits: Grokking as Competition of Sparse and Dense Subnetworks, March 2023
Merrill, W., Tsilivis, N., and Shukla, A · 2023
Closest in time.
The On-Line Encyclopedia of Integer Sequences, 2023
OEIS Foundation Inc · 2023
Closest in time.
Rubin, N., Seroussi, I., and Ringel, Z · 2023
Closest in time.
Sage Mathematics Software (Version 10.0.0)
Stein, W. et al · 2023
Closest in time.
Linear representations of sentiment in large language models, 2023
Tigges, C., Hollinsworth, O. J., Geiger, A., and Nanda, N · 2023
Closest in time.
Explaining grokking through circuit efficiency, September 2023
Varma, V., Shah, R., Kenton, Z., Kramár, J., and Kumar, R · 2023
Closest in time.
pola-rs/polars: Python Polars 0.19.0, August 2023
Vink, R., Gooijer, S. d., Beedie, A., Gorelli, M. E., Zundert, J. v., Hulselmans, G., Grinstead, C., Santamaria, M., Guo, W., Heres, D., Magarick, J., Marshall, ibENPC, Peters, O., Leitao, J., Wilksch, M., Heerden, M. v., Borchert, O., Jermain, C., Haag, J., Peek, J., Russell, R., Pryer, C., Castellanos, A. G., Goh, J., illumination-k, Brannigan, L., Conradt, M., and Robert · 2023
Closest in time.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Wen, K., Li, Y., Liu, B., and Risteski, A · 2023
Closest in time.
Benign Overfitting and Grokking in ReLU Networks for XOR Cluster Data, October 2023
Xu, Z., Wang, Y., Frei, S., Vardi, G., and Hu, W · 2023
Closest in time.
Can Transformers Learn to Solve Problems Recursively?, June 2023
Zhang, S. D., Tigges, C., Biderman, S., Raginsky, M., and Ringer, T · 2023
Closest in time.
The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks, June 2023
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Closest in time.
Feature emergence via margin maximization: case studies in algebraic tasks, 2024
Morwani, D., Edelman, B. L., Oncescu, C.-A., Zhao, R., and Kakade, S · 2024
Closest in time.
Understanding addition in transformers, 2024
Quirke, P. and Barez, F · 2024
Closest in time.