Fetching the paper…
Reading the bibliography…
Training Large Language Models (LLMs) presents significant memory challenges, predominantly due to the growing size of weights and optimizer states.
Chen, Y. and Wainwright, M. J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Training Deep Nets with Sublinear Memory Cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Gradient Descent Happens in a Tiny Subspace
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Earlier work this paper cites.
Gradient-based meta-learning with learned layerwise metric and subspace
Lee, Y. and Choi, S · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Atomo: Communication-efficient learning via atomic sparsification
Wang, H., Sievert, S., Liu, S., Charles, Z., Papailiopoulos, D., and Wright, S · 2018
Earlier work this paper cites.
Inrank: Incremental low-rank learning
Zhao, J., Zhang, Y., Chen, B., Schäfer, F., and Anandkumar, A · 2018
Earlier work this paper cites.
Memory efficient adaptive optimization
Anil, R., Gupta, V., Koren, T., and Singer, Y · 2019
Earlier work this paper cites.
Non-Convex Projected Gradient Descent for Generalized Low-Rank Tensor Regression
Chen, H., Raskutti, G., and Yuan, M · 2019
Earlier work this paper cites.
Dynamic mini-batch sgd for elastic distributed training: Learning in the limbo of resources
Lin, H., Zhang, H., Ma, Y., He, T., Zhang, Z., Zha, S., and Li, M · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Continual learning in low-rank orthogonal subspaces
Chaudhry, A., Khan, N., Dokania, P., and Torr, P · 2020
Earlier work this paper cites.
Low-rank gradient approximation for memory-efficient on-device training of deep neural network
Gooneratne, M., Sim, K. C., Zadrazil, P., Kabel, A., Beaufays, F., and Motta, G · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Glu variants improve transformer
Shazeer, N · 2020
Cited alongside, same era.
Understanding self-supervised learning with dual deep networks
Tian, Y., Yu, L., Chen, X., and Ganguli, S · 2020
Cited alongside, same era.
Low-Rank Gradient Descent
Cosson, R., Jadbabaie, A., Makur, A., Reisizadeh, A., and Shah, D · 2023
Later among the works it cites.
Low-Rank Gradient Descent for Memory-Efficient Training of Deep In-Memory Arrays
Huang, S., Hoskins, B. D., Daniels, M. W., Stiles, M. D., and Adam, G. C · 2023
Later among the works it cites.
Error Feedback Can Accurately Compress Preconditioners
Modoranu, I.-V., Kalinov, A., Kurtic, E., Frantar, E., and Alistarh, D · 2023
Later among the works it cites.
Tied-Lora: Enhacing parameter efficiency of LoRA with weight tying
Renduchintala, A., Konuk, T., and Kuchaiev, O · 2023
Later among the works it cites.
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., Gonzalez, J. E., and Stoica, I · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Practical low-rank communication compression in decentralized deep learning
Vogels, T., Karimireddy, S. P., and Jaggi, M · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Cited alongside, same era.
8-bit optimizers via block-wise quantization
Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L · 2022
Cited alongside, same era.
Delta Tuning: A Comprehensive Study of Parameter Efficient Methods for Pre-trained Language Models
Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., Yi, J., Zhao, W., Wang, X., Liu, Z., Zheng, H.-T., Chen, J., Liu, Y., Tang, J., Li, J., and Sun, M · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Exploring Low Rank Training of Deep Neural Networks
Kamalakara, S. R., Locatelli, A., Venkitesh, B., Ba, J., Gal, Y., and Gomez, A. N · 2022
Cited alongside, same era.
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Stable and low-precision training for large-scale vision-language models
Wortsman, M., Dettmers, T., Zettlemoyer, L., Morcos, A., Farhadi, A., and Schmidt, L · 2023
Later among the works it cites.
A spectral condition for feature learning
Yang, G., Simon, J. B., and Bernstein, J · 2023
Later among the works it cites.
Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning
Zhang, L., Zhang, L., Shi, S., Chu, X., and Li, B · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
Flora: Low-Rank Adapters Are Secretly Gradient Compressors
Hao, Y., Cao, Y., and Mou, L · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2024
Closest in time.
Memory efficient optimizers with 4-bit states
Li, B., Chen, J., and Zhu, J · 2024
Closest in time.
ReloRA: High-rank training through low-rank updates
Lialin, V., Muckatira, S., Shivagunde, N., and Rumshisky, A · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
JoMA: Demystifying multilayer transformers via joint dynamics of MLP and attention
Tian, Y., Wang, Y., Zhang, Z., Chen, B., and Du, S. S · 2024
Closest in time.
Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning
Xia, W., Qin, C., and Hazan, E · 2024
Closest in time.