Fetching the paper…
Reading the bibliography…
Compression has been a critical lens to understand the success of Transformers.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2005
Earlier work this paper cites.
The hutter prize
Marcus Hutter · 2006
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Earlier work this paper cites.
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli · 2019
Earlier work this paper cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André FT Martins · 2019
Earlier work this paper cites.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Sam Smith · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2020
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, et al · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
White-box transformers via sparse rate reduction
Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin Haeffele, and Yi Ma · 2023
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2023
Later among the works it cites.
A survey on mixture of experts
Weilin Cai, Juyong Jiang, Fan Wang, Jing Tang, Sunghun Kim, and Jiayi Huang · 2024
Later among the works it cites.
Understanding emergent abilities of language models from the loss perspective
Zhengxiao Du, Aohan Zeng, Yuxiao Dong, and Jie Tang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Cited alongside, same era.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Cited alongside, same era.
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al · 2023
Cited alongside, same era.
Simplifying transformer blocks
Bobby He and Thomas Hofmann · 2023
Cited alongside, same era.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L Smith, and Yee Whye Teh · 2023
Cited alongside, same era.
Same pre-training loss, better downstream: Implicit bias matters for language models
Hong Liu, Sang Michael Xie, Zhiyuan Li, and Tengyu Ma · 2023
Cited alongside, same era.
A theory on adam instability in large-scale machine learning
Igor Molybog, Peter Albert, Moya Chen, Zachary DeVito, David Esiobu, Naman Goyal, Punit Singh Koura, Sharan Narang, Andrew Poulton, Ruan Silva, et al · 2023
Cited alongside, same era.
Yuzhen Huang, Jinghan Zhang, Zifei Shan, and Junxian He · 2024
Later among the works it cites.
Scaling laws for fine-grained mixture of experts
Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski, Maciej Pióro, Michał Krutul, Szymon Antoniak, Kamil Ciebiera, Krystian Król, Tomasz Odrzygóźdź, Piotr Sankowski, et al · 2024
Later among the works it cites.
Dynamic sparsity in the brain and machines routing information through neural pathways
André Martins, Edoardo Ponti, Duarte Alves, Piotr Nawrot, and Saul Santos · 2024
Later among the works it cites.
Transformers are multi-state rnns
Matanel Oren, Michael Hassid, Nir Yarden, Yossi Adi, and Roy Schwartz · 2024
Later among the works it cites.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro · 2024
Later among the works it cites.
Implicit regularization of gradient flow on one-layer softmax attention
Heejune Sheen, Siyu Chen, Tianhao Wang, and Harrison H Zhou · 2024
Later among the works it cites.
Implicit bias and fast convergence rates for self-attention
Bhavya Vasudeva, Puneesh Deora, and Christos Thrampoulidis · 2024
Later among the works it cites.
Scaling white-box transformers for vision
Jinrui Yang, Xianhang Li, Druv Pai, Yuyin Zhou, Yi Ma, Yaodong Yu, and Cihang Xie · 2024
Later among the works it cites.
Understanding llm behaviors via compression: Data generation, knowledge acquisition and scaling laws
Zhixuan Pan, Shaowen Wang, and Jian Li · 2025
Closest in time.
Attention-only transformers via unrolled subspace denoising
Peng Wang, Yifu Lu, Yaodong Yu, Druv Pai, Qing Qu, and Yi Ma · 2025
Closest in time.