Fetching the paper…
Reading the bibliography…
As language models grow ever larger, so do their vocabularies.
Megatron-LM: Training multi-billion parameter language models using model parallelism, 2019
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 1909
Earlier work this paper cites.
Pracniques: further remarks on reducing truncation errors
William Kahan · 1965
Earlier work this paper cites.
Data parallel algorithms
W. Daniel Hillis and Guy L. Steele · 1986
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Linformer: Self-attention with linear complexity, 2020
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour, 2017
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2017
Earlier work this paper cites.
CUTLASS: Fast linear algebra in CUDA C++, 2017
Andrew Kerr, Duane Merrill, Julien Demouth, and John Tran · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Online normalizer calculation for softmax, 2018
Maxim Milakov and Natalia Gimelshein · 2018
Earlier work this paper cites.
Openwebtext corpus, 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Pipedream: Generalized pipeline parallelism for DNN training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Triton: An intermediate language and compiler for tiled neural network computations
Philippe Tillet, Hsiang-Tsung Kung, and David D. Cox · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing, 2019
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Sequence parallelism: Long sequence training from system perspective
Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You · 2023
Later among the works it cites.
Stanford Alpaca: An instruction-following LLaMA model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
MEGABYTE: Predicting million-byte sequences with multiscale transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Marah I Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat S. Behl, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
ZeRO: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller · 2021
Cited alongside, same era.
Self-attention does not need O(n 2 {}^{\mbox{2}} ) memory, 2021
Markus N. Rabe and Charles Staats · 2021
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
Speed is all you need: On-device acceleration of large diffusion models via GPU-aware optimizations
Yu-Hui Chen, Raman Sarokin, Juhyun Lee, Jiuqiang Tang, Chuo-Ling Chang, Andrei Kulik, and Matthias Grundmann · 2023
Cited alongside, same era.
Closest in time.
Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation
Jason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, et al · 2024
Closest in time.
The Llama 3 herd of models, 2024
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Liger-Kernel: Efficient Triton kernels for LLM training, 2024
Pin-Lun Hsu, Yun Dai, Vignesh Kothapalli, Qingquan Song, Shao Tang, and Siyu Zhu · 2024
Closest in time.
Mistral NeMo, 2024
Mistral AI Team · 2024
Closest in time.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Morgane Rivière, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, et al · 2024
Closest in time.
Scaling laws with vocabulary: Larger models deserve larger vocabularies, 2024
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong · 2024
Closest in time.
torchtune, 2024
Torch Tune Team · 2024
Closest in time.