Fetching the paper…
Reading the bibliography…
We present the first comprehensive study of latent multi-head attention (MLA) for small language models, revealing interesting efficiency-quality trade-offs.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
2006
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems 30
2017
Earlier work this paper cites.
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-Attention with Relative Position Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2018)
2018
Earlier work this paper cites.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Pranav Shyam, Girish Sastry, Amanda Askell, and others. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems 33
2020
Earlier work this paper cites.
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The Efficient Transformer. In International Conference on Learning Representations (ICLR 2020)
2020
Earlier work this paper cites.
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A Lite BERT for Self- Supervised Learning of Language Representations. In International Conference on Learning Representations (ICLR 2020)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, and others. 2021. Rethinking Attention with Performers. In International Conference on Learning Representations (ICLR 2021)
2021
Earlier work this paper cites.
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu. 2021. RoFormer: Enhanced Transformer with Rotary Position Embedding. In Proceedings of ACL-IJCNLP 2021
2021
Cited alongside, same era.
Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. (2021)
2021
Cited alongside, same era.
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Advances in Neural Information Processing Systems 35
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774
2023
Cited alongside, same era.
2024
Later among the works it cites.
MobiLLM Team. 2024. MobiLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases. In International Conference on Machine Learning (ICML 2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.