Fetching the paper…
Reading the bibliography…
Training LLMs presents significant memory challenges due to growing size of data, weights, and optimizer states.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
New perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
ATOMO: Communication-efficient learning via atomic sparsification
Shiqiang Wang, Gauri Joshi, Sreeram K Ghosh, and H Vincent Poor · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2018
Earlier work this paper cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Menglong Chen, Denny Chen, Zhifeng Hu, Yuxin Shen, Maxim Krikun, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Rohan Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Low-rank gradient approximation for multi-task learning
Shamal Gooneratne, Meng Wang, Zhili Guo, Vamsi Krishna Kanuparthi, Dinesh Rajan, and Anura P Jayasumana · 2020
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
ZeRO: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
ZerO initialization: Initializing neural networks with zero-valued parameters
Shangqian Zhao, Shiyu Li, and Yi Ma · 2022
Later among the works it cites.
Low-rank gradient descent converges and generalizes
Victor Cosson, Baptiste Lecouat, Arthur Varre, Stéphane d’Ascoli, and Giulio Biroli · 2023
Later among the works it cites.
QLoRA: Efficient finetuning of quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Mistral: Efficient composable inference for large language models
Ye Jiang, Pengcheng Li, Zhe Gan, Jianfeng Liu, Dongdong Chen, Xiaodong Zhu, Zhangyang Li, Lijuan Wang, Jianfeng Wang, and Zicheng Liu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Noam Shazeer · 2020
Cited alongside, same era.
PowerGossip: Practical low-rank communication for decentralized optimization
Thijs Vogels, Martin Jaggi, and Giorgio Patrini · 2020
Cited alongside, same era.
Extending torchelastic for stateful training jobs
Tianshi Zhao, Zhen Sun, Xiaodong Wang, Fei Zhou, Yang Guo, and Alexander J Smola · 2020
Cited alongside, same era.
Pengcheng He, Jianfeng Gao, and Weizhu Chen · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ning Ding, Xiang Zheng, Yujia Wang, Yifei Chen, Yichi Liu, Haitao Zheng, Xipeng Qiu, Yujun Shen, Bolin Ding, and Jie Tang · 2022
Cited alongside, same era.
Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W Mahoney, Kurt Keutzer, and Amir Gholami · 2023
Later among the works it cites.
ReLoRA: Low-rank fine-tuning reloaded
Vladimir Lialin and Arthur Schatz · 2023
Later among the works it cites.
Tied lora: Enhancing parameter-efficient fine-tuning with tied weights
Adithya Renduchintala, Pedro Rodriguez, and Mathias Creutz · 2023
Later among the works it cites.
S-LoRA: Scalable efficient model serving for massive lora models
Yi Sheng, Xuefei Han, Xuefeng Zhu, Yuanzhi Yang, Jiani Sun, and Guohui Zhou · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Multi-LoRA: Efficient fine-tuning for democratic AI
Zihao Wang, Zhen Bai, and Sophia Ananiadou · 2023
Later among the works it cites.
Spectral methods in low-rank model adaptation
Zhilin Yang, Edward J Hu, Tianle Xia, Richard Socher, and Yuanzhi Li · 2023
Later among the works it cites.
LoRA-FA: Memory-efficient low-rank adaptation via feature re-alignment
Rui Zhang et al · 2023
Later among the works it cites.
TinyAgent: Function calling at the edge
Lutfi Eren Erdogan, Nicholas Lee, Siddharth Jha, Sehoon Kim, Ryan Tabrizi, Suhong Moon, Coleman Hooper, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami · 2024
Closest in time.
Chain-of-thought lora: Efficient adaptation of large language models
Tianxiang Xia, Hao Peng, Zheyu Chen, Lemao Li, Zhiyuan He, Zhen Yang, and Wei-Ying Ma · 2024
Closest in time.