Fetching the paper…
Reading the bibliography…
Understanding the learning process and the embedded computation in transformers is becoming a central goal for the development of interpretable AI.
Trainable grammars for speech recognition
Baker, J. K · 1979
Earlier work this paper cites.
Inside-outside probability computation for belief propagation
Sato, T · 2007
Earlier work this paper cites.
Information, physics, and computation
Mézard, M. and Montanari, A · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Belief propagation, robust reconstruction and optimal recovery of block models
Mossel, E., Neeman, J., and Sly, A · 2014
Earlier work this paper cites.
Deep learning and hierarchal generative models
Mossel, E · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Random language model
De Giuli, E · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Curriculum learning for language modeling
Campos, D · 2021
Cited alongside, same era.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Cited alongside, same era.
Physics of language models: Part 1, context-free grammar
Allen-Zhu, Z. and Li, Y · 2023
Cited alongside, same era.
Do transformers parse while predicting the masked word?
Zhao, H., Panigrahi, A., Ge, R., and Arora, S · 2023
Later among the works it cites.
Sliding down the stairs: How correlated latent variables accelerate learning with neural networks
Bardone, L. and Goldt, S · 2024
Closest in time.
Understanding counting in small transformers: The interplay between attention and feed-forward layers
Behrens, F., Biggio, L., and Zdeborová, L · 2024
Closest in time.
Towards a theory of how the structure of language is acquired by deep neural networks
Cagnetta, F. and Wyart, M · 2024
Closest in time.
How deep neural networks learn compositional data: The random hierarchy model
Cagnetta, F., Petrini, L., Tomasini, U. M., Favero, A., and Wyart, M · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Approximating CKY with transformers
Khalighinejad, G., Liu, O., and Wiseman, S · 2023
Cited alongside, same era.
Exact phase transitions for stochastic block models and reconstruction on trees
Mossel, E., Sly, A., and Sohn, Y · 2023
Cited alongside, same era.
Neural networks trained with SGD learn distributions of increasing complexity
Refinetti, M., Ingrosso, A., and Goldt, S · 2023
Cited alongside, same era.
Motif learning facilitates sequence memorization and generalization
Wu, S., Thalmann, M., and Schulz, E · 2023
Cited alongside, same era.
Applications of transformer-based language models in bioinformatics: a survey
Zhang, S., Fan, R., Liu, Y., Chen, S., Liu, Q., and Zeng, W · 2023
Cited alongside, same era.
Mei, S · 2024
Closest in time.
Tulip: A transformer-based unsupervised language model for interacting peptides and t cell receptors that generalizes to unseen epitopes
Meynard-Piganeau, B., Feinauer, C., Weigt, M., Walczak, A. M., and Mora, T · 2024
Closest in time.
A distributional simplicity bias in the learning dynamics of transformers
Rende, R., Gerace, F., Laio, A., Goldt, S., et al · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2024
Closest in time.
A phase transition in diffusion models reveals the hierarchical nature of data
Sclocchi, A., Favero, A., and Wyart, M · 2025
Closest in time.