Fetching the paper…
Reading the bibliography…
Padding is often used in tuning LLM models by adding special tokens to shorter training examples to match the length of the longest sequence in each batch.
The stack: 3 tb of permissively licensed source code
Denis Kocetkov et al · 2022
Earlier work this paper cites.
Finetuned language models are zero-shot learners, 2022
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Tri Dao · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
Efficient sequence packing without cross-contamination: Accelerating large language models without impacting performance, 2023
Mario Michael Krell, Matej Kosec, Sergio P. Perez, and Andrew William Fitzgibbon · 2023
Cited alongside, same era.
Code llama: Open foundation models for code, 2023
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei · 2024
Cited alongside, same era.
URL https://github.com/Dao-AILab/flash-attention/issues/654
Flash attention issue 654
Cited in the paper.
Saving memory using padding-free transformer layers during finetuning, 2024
Mayank Mishra · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math, 2024
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah · 2024
Closest in time.
Analysing the impact of sequence composition on language model pre-training
Yu Zhao, Yuanbin Qu, Konrad Staniszewski, Szymon Tworkowski, Wei Liu, Piotr Miłoś, Yuxiang Wu, and Pasquale Minervini · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
URL hhttps://huggingface.co/blog/packing-with-FA2
Improving hugging face training efficiency through packing with flash attention
Cited in the paper.
URL https://github.com/huggingface/transformers/pull/31629
Enhancing sft training efficiency using packing and flashattention2 with position ids
Cited in the paper.
URL https://github.com/OpenAccess-AI-Collective/axolotl/tree/main/src/axolotl/monkeypatch
Axolotl
Cited in the paper.
URL https://huggingface.co/transformers/v4.4.2/_modules/transformers/trainer_pt_utils.html
Lengthgroupedsampler in hugging face transformers library
Cited in the paper.
URL https://github.com/imoneoi/multipack_sampler/tree/master
Multi-pack sampler repository
Cited in the paper.
URL https://github.com/Sanster/padding_free_llm_train/tree/main
Padding free llm train
Cited in the paper.