Fetching the paper…
Reading the bibliography…
In this work, we study rapid improvements of the training loss in transformers when being confronted with multi-step decision tasks.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Earlier work this paper cites.
Guiding attention for self-supervised learning with transformers
Deshpande, A. and Narasimhan, K · 2020
Earlier work this paper cites.
Increasing learning efficiency of self-attention networks through direct position interactions, learnable temperature, and convoluted attention
Dufter, P., Schmitt, M., and Schütze, H · 2020
Earlier work this paper cites.
Gmat: Global memory augmentation for transformers
Gupta, A. and Berant, J · 2020
Earlier work this paper cites.
Improving transformer optimization through better initialization
Huang, X. S., Perez, F., Ba, J., and Volkovs, M · 2020
Earlier work this paper cites.
Contrastive multiview coding
Tian, Y., Krishnan, D., and Isola, P · 2020
Cited alongside, same era.
Information-theoretic probing with minimum description length
Voita, E. and Titov, I · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Xcit: Cross-covariance image transformers
Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Thilak, V., Littwin, E., Zhai, S., Saremi, O., Paiss, R., and Susskind, J. M · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Later among the works it cites.
Accumulated trivial attention matters in vision transformers on small datasets
Chen, X., Hu, Q., Li, K., Zhong, C., and Wang, G · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
Intriguing Properties of Transformer Training Instabilities, 2023
Gilmer, J., Schioppa, A., and Cohen, J · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H · 2021
Cited alongside, same era.
Derivative of the softmax function and the categorical cross-entropy loss
Kurbiel, T · 2021
Cited alongside, same era.
Pre-training a bert with curriculum learning by increasing block-size of input text
Nagatsuka, K., Broni-Bediako, C., and Atsumi, M · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H · 2021
Cited alongside, same era.
Escaping the gradient vanishing: Periodic alternatives of softmax in attention mechanism
Wang, S., Liu, F., and Liu, B · 2021
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Barak, B., Edelman, B., Goel, S., Kakade, S., Malach, E., and Zhang, C · 2022
Cited alongside, same era.
Data distributional properties drive emergent in-context learning in transformers
Chan, S., Santoro, A., Lampinen, A., Wang, J., Singh, A., Richemond, P., McClelland, J., and Hill, F · 2022
Cited alongside, same era.
Normsoftmax: Normalizing the input of softmax to accelerate and stabilize training
Jiang, Z., Gu, J., and Pan, D. Z · 2023
Closest in time.
Omnigrok: Grokking beyond algorithmic data
Liu, Z., Michaud, E. J., and Tegmark, M · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Liberum, T., Smith, J., and Steinhardt, J · 2023
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Reddy, G · 2023
Closest in time.
A study on relu and softmax in transformer
Shen, K., Guo, J., Tan, X., Tang, S., Wang, R., and Bian, J · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Closest in time.
Stabilizing transformer training by preventing attention entropy collapse
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Chen, A., Shwartz-Ziv, R., Cho, K., Leavitt, M. L., and Saphra, N · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Closest in time.
Future ml systems will be qualitatively different
Steinhardt, J · 2024
Closest in time.