Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have revolutionized natural language understanding and generation but face significant memory bottlenecks during training.
Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions
Nathan Halko, Per-Gunnar Martinsson, and Joel A Tropp · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python, 2018
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Andreas Müller, Joel Nothman, Gilles Louppe, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
Tensorly: Tensor learning in python
Jean Kossaifi, Yannis Panagakis, Anima Anandkumar, and Maja Pantic · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Conqur: Mitigating delusional bias in deep q-learning
Dijia Su, Jayden Ooi, Tyler Lu, Dale Schuurmans, and Craig Boutilier · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
DiJia Su, Jason D Lee, John M Mulvey, and H Vincent Poor · 2021
Earlier work this paper cites.
8-bit optimizers via block-wise quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Narrowing the coordinate-frame gap in behavior prediction models: Distillation for efficient and accurate scene-centric motion forecasting
DiJia Andy Su, Bertrand Douillard, Rami Al-Rfou, Cheol Park, and Benjamin Sapp · 2022
Earlier work this paper cites.
Relora: High-rank training through low-rank updates
Vladislav Lialin, Sherin Muckatira, Namrata Shivagunde, and Anna Rumshisky · 2023
Earlier work this paper cites.
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
Kai Lv, Hang Yan, Qipeng Guo, Haijun Lv, and Xipeng Qiu · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Multilora: Democratizing lora for better multi-task learning
Yiming Wang, Yu Lin, Xiaodong Zeng, and Guannan Zhang · 2023
Cited alongside, same era.
Natural galore: Accelerating galore for memory-efficient llm training and fine-tuning
Arijit Das · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Later among the works it cites.
Ldadam: Adaptive optimization from low-dimensional gradient statistics
Thomas Robert, M. H. Safaryan, Ionut-Vlad Modoranu, and Dan Alistarh · 2024
Later among the works it cites.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
DiJia Su, Sainbayar Sukhbaatar, Michael Rabbat, Yuandong Tian, and Qinqing Zheng · 2024
Later among the works it cites.
Learning personalized alignment for evaluating open-ended text generation
Danqing Wang, Kevin Yang, Hanlin Zhu, Xiaomeng Yang, Andrew Cohen, Lei Li, and Yuandong Tian · 2024
Later among the works it cites.
Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert Joseph George, David Pitt, Jiawei Zhao, Jean Kossaifi, Cheng Luo, Yuandong Tian, and Anima Anandkumar · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Cited alongside, same era.
From galore to welore: How low-rank weights non-uniformly emerge from low-rank gradients
Ajay Kumar Jaiswal, Lu Yin, Zhenyu (Allen) Zhang, Shiwei Liu, Jiawei Zhao, Yuandong Tian, and Zhangyang Wang · 2024
Cited alongside, same era.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lucas Lehnert, Sainbayar Sukhbaatar, DiJia Su, Qinqing Zheng, Paul Mcvay, Michael Rabbat, and Yuandong Tian · 2024
Cited alongside, same era.
Memory-efficient llm training with online subspace descent, 2024
Kaizhao Liang, Bo Liu, Lizhang Chen, and Qiang Liu · 2024
Cited alongside, same era.
Badam: A memory efficient full parameter training method for large language models
Qijun Luo, Hengxu Yu, and Xiao Li · 2024
Cited alongside, same era.
Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning
Rui Pan, Xiang Liu, Shizhe Diao, Renjie Pi, Jipeng Zhang, Chi Han, and Tong Zhang · 2024
Cited alongside, same era.
Spinquant: Llm quantization with learned rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort
Cited in the paper.
Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar · 2024
Later among the works it cites.
Q-galore: Quantized galore with int4 projection and layer-adaptive low-rank gradients
Zhenyu Zhang, Ajay Jaiswal, Lu Yin, Shiwei Liu, Jiawei Zhao, Yuandong Tian, and Zhangyang Wang · 2024
Later among the works it cites.
Galore: Memory-efficient llm training by gradient low-rank projection
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian · 2024
Later among the works it cites.
Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks
Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li · 2024
Later among the works it cites.
Spectral journey: How transformers predict the shortest path
Andrew Cohen, Andrey Gromov, Kaiyu Yang, and Yuandong Tian · 2025
Closest in time.
Step-kto: Optimizing mathematical reasoning through stepwise binary feedback
Yen-Ting Lin, Di Jin, Tengyu Xu, Tianhao Wu, Sainbayar Sukhbaatar, Chen Zhu, Yun He, Yun-Nung Chen, Jason Weston, Yuandong Tian, et al · 2025
Closest in time.
Token assorted: Mixing latent and text tokens for improved language model reasoning
DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng · 2025
Closest in time.