Fetching the paper…
Reading the bibliography…
We introduce LDAdam, a memory-efficient optimizer for training large models, that performs adaptive optimization steps within lower dimensional subspaces, while consistently exploring the full parameter space during training.
Error feedback fixes signsgd and other gradient compression schemes, 2019
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 1901
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp, 2019
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 1902
Earlier work this paper cites.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 1904
Earlier work this paper cites.
Powersgd: Practical low-rank gradient compression for distributed optimization, 2020
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Root mean square layer normalization, 2019
Biao Zhang and Rico Sennrich · 1910
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2002
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Parameter-efficient transfer learning with diff pruning, 2021
Demi Guo, Alexander M. Rush, and Yoon Kim · 2012
Earlier work this paper cites.
Lecture 6e rmsprop: Divide the gradient by a running average of its recent magnitude
Geoffrey Hinton · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Block Power Method for SVD Decomposition
Abdeslem Bentbib and A. Kanber · 2015
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding, 2017
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Implicit regularization in deep learning, 2017
Behnam Neyshabur · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization, 2017
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods, 2018
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes, 2018
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost, 2018
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Convergence of adam for non-convex objectives: Relaxed hyperparameters and non-ergodic case
Meixuan He, Yuqing Liang, Jinlan Liu, and Dongpo Xu · 2023
Later among the works it cites.
Convergence of Adam under relaxed assumptions
Haochuan Li, Ali Jadbabaie, and Alexander Rakhlin · 2023
Later among the works it cites.
Relora: High-rank training through low-rank updates, 2023
Vladislav Lialin, Namrata Shivagunde, Sherin Muckatira, and Anna Rumshisky · 2023
Later among the works it cites.
Came: Confidence-guided adaptive memory efficient optimization, 2023
Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang, and Yang You · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training, 2020
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally · 2020
Cited alongside, same era.
Rmsprop converges with proper hyperparameter
Naichen Shi, Dawei Li, Mingyi Hong, and Ruoyu Sun · 2020
Cited alongside, same era.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Adalora: Adaptive budget allocation for parameter-efficient fine-tuning, 2023
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao · 2023
Later among the works it cites.
Parameter-efficient fine-tuning for large models: A comprehensive survey, 2024
Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang · 2024
Closest in time.
On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
Yusu Hong and Junhong Lin · 2024
Closest in time.
From galore to welore: How low-rank weights non-uniformly emerge from low-rank gradients, 2024
Ajay Jaiswal, Lu Yin, Zhenyu Zhang, Shiwei Liu, Jiawei Zhao, Yuandong Tian, and Zhangyang Wang · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation, 2024
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen · 2024
Closest in time.
Microadam: Accurate adaptive optimization with low space overhead and provable convergence, 2024
Ionut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic, Thomas Robert, Peter Richtarik, and Dan Alistarh · 2024
Closest in time.
Rosa: Accurate parameter-efficient fine-tuning via robust adaptation, 2024
Mahdi Nikdan, Soroush Tabesh, Elvir Crnčević, and Dan Alistarh · 2024
Closest in time.
How to save memory by fusing the optimizer step into the backward pass, n.d
Pytorch_Tutorials · 2024
Closest in time.
ADOPT: Modified Adam Can Converge with Any β 2 \beta_{2} with the Optimal Rate
Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima, Seong Cheol Jeong, Go Nagahara, Tomoshi Iiyama, Masahiro Suzuki, Yusuke Iwasawa, and Yutaka Matsuo · 2024
Closest in time.
Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhi-Ming Ma, Tie-Yan Liu, and Wei Chen · 2024
Closest in time.
Chain of lora: Efficient fine-tuning of language models via residual learning, 2024
Wenhan Xia, Chengwei Qin, and Elad Hazan · 2024
Closest in time.
Galore: Memory-efficient llm training by gradient low-rank projection, 2024
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian · 2024
Closest in time.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Jinghui Chen, Yuan Cao, Ziyan Yang, and Quanquan Gu · 2024
Closest in time.