Fetching the paper…
Reading the bibliography…
Next to scaling considerations, architectural design choices profoundly shape the solution space of transformers.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 1912
Earlier work this paper cites.
Lower bounds on the maximum cross correlation of signals (corresp.)
Welch, L. R · 1974
Earlier work this paper cites.
Optimally sparse representation in general (nonorthogonal) dictionaries via sup minimization
Donoho, D. L. and Elad, M · 2003
Earlier work this paper cites.
Grassmannian frames with applications to coding and communication
Strohmer, T. and Heath, R. W · 2003
Earlier work this paper cites.
Causal mediation analysis for interpreting neural NLP: the case of gender bias
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. M · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Tables of the existence of equiangular tight frames, 2016
Fickus, M. and Mixon, D. G · 2016
Earlier work this paper cites.
Gradient-based algorithm for designing sensing matrix considering real mutual coherence for compressed sensing systems
Jiang, Q., Li, S., Bai, H., de Lamare, R. C., and He, X · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Video action transformer network
Girdhar, R., Carreira, J., Doersch, C., and Zisserman, A · 2019
Earlier work this paper cites.
Thread: circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Earlier work this paper cites.
Abstraction and reasoning challenge, 2020
Chollet, F., Tong, K., Reade, W., and Elliott, J · 2020
Earlier work this paper cites.
interpreting GPT: the logit lens — LessWrong, January 2020
nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Notes on the symmetries of 2-layer relu-networks
Petzka, H., Trimmel, M., and Sminchisescu, C · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M. I., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Cited alongside, same era.
MLP-mixer: An all-MLP architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A. P., Keysers, D., Uszkoreit, J., Lucic, M., and Dosovitskiy, A · 2021
Cited alongside, same era.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Cited alongside, same era.
Toy models of superposition, 2022
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Later among the works it cites.
The quantization model of neural scaling
Michaud, E. J., Liu, Z., Girit, U., and Tegmark, M · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Counting and algorithmic generalization with transformers
Ouellette, S., Pfister, R., and Jud, H · 2023
Later among the works it cites.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventos, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Telet: A monotonic algorithm to design large dimensional equiangular tight frames for applications in compressed sensing
Jyothi, R. and Babu, P · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A. J., and Belinkov, Y · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Cited alongside, same era.
Generalization on the unseen, logic reasoning and degree curriculum
Abbe, E., Bengio, S., Lotfi, A., and Rizk, K · 2023
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Delétang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., and Ortega, P. A · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., Hwang, J., Sanyal, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2023
Cited alongside, same era.
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Later among the works it cites.
Summing up the facts: Additive mechanisms behind factual recall in LLMs, 2024
Chughtai, B., Cooney, A., and Nanda, N · 2024
Closest in time.
A phase transition between positional and semantic learning in a solvable model of dot-product attention
Cui, H., Behrens, F., Krzakala, F., and Zdeborova, L · 2024
Closest in time.
Rethinking attention: Exploring shallow feed-forward neural networks as an alternative to attention layers in transformers (student abstract)
Dordevic, D., Bozic, V., Thommes, J., Coppola, D., and Pal Singh, S · 2024
Closest in time.
Contextual counting: A mechanistic study of transformers on a quantitative task, 2024
Golkar, S., Bietti, A., Pettee, M., Eickenberg, M., Cranmer, M., Hirashima, K., Krawezik, G., Lourie, N., McCabe, M., Morel, R., Ohana, R., Parker, L. H., Blancard, B. R.-S., Cho, K., and Ho, S · 2024
Closest in time.
The llama 3 herd of models, 2024
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., and et al · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Closest in time.
Understanding addition in transformers
Quirke, P. and Barez, F · 2024
Closest in time.
When can transformers count to n?, 2024
Yehudai, G., Kaplan, H., Ghandeharioun, A., Geva, M., and Globerson, A · 2024
Closest in time.
Not all language model features are one-dimensionally linear
Engels, J., Michaud, E. J., Liao, I., Gurnee, W., and Tegmark, M · 2025
Closest in time.
Understanding factual recall in transformers via associative memories
Nichani, E., Lee, J. D., and Bietti, A · 2025
Closest in time.