Fetching the paper…
Reading the bibliography…
Parameter-efficient fine-tuning (PEFT) significantly reduces memory costs when adapting large language models (LLMs) for downstream applications.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C Spall · 1992
Earlier work this paper cites.
Backpropagation and stochastic gradient descent method
Shun-ichi Amari · 1993
Earlier work this paper cites.
Building a question answering test collection
Ellen M Voorhees and Dawn M Tice · 2000
Earlier work this paper cites.
Convergence of a block coordinate descent method for nondifferentiable minimization
Paul Tseng · 2001
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Coordinate descent algorithms
Stephen J Wright · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order
Xiangru Lian, Huan Zhang, Cho-Jui Hsieh, Yijun Huang, and Ji Liu · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Iteration complexity analysis of block coordinate descent methods
Mingyi Hong, Xiangfeng Wang, Meisam Razaviyayn, and Zhi-Quan Luo · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Cited alongside, same era.
Black-box adversarial attacks with limited queries and information
Andrew Ilyas, Logan Engstrom, Anish Athalye, and Jessy Lin · 2018
Cited alongside, same era.
Zeroth-order stochastic variance reduction for nonconvex optimization
Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Paishun Ting, Shiyu Chang, and Lisa Amini · 2018
Cited alongside, same era.
Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates
Krishnakumar Balasubramanian and Saeed Ghadimi · 2018
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Later among the works it cites.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Sijia Liu, Pin-Yu Chen, Bhavya Kailkhura, Gaoyuan Zhang, Alfred O Hero III, and Pramod K Varshney · 2020
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Cited alongside, same era.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Cited alongside, same era.
Record: Bridging the gap between human and machine commonsense reading comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme · 2018
Cited alongside, same era.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Cited alongside, same era.
Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization
Xiangyi Chen, Sijia Liu, Kaidi Xu, Xingguo Li, Xue Lin, Mingyi Hong, and David Cox · 2019
Cited alongside, same era.
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
How many degrees of freedom do we need to train deep networks: a loss landscape perspective
Brett W Larsen, Stanislav Fort, Nic Becker, and Surya Ganguli · 2021
Later among the works it cites.
A theoretical and empirical comparison of gradient approximations in derivative-free optimization
Albert S Berahas, Liyuan Cao, Krzysztof Choromanski, and Katya Scheinberg · 2022
Later among the works it cites.
Zarts: On zero-order optimization for neural architecture search
Xiaoxing Wang, Wenxuan Guo, Jianlin Su, Xiaokang Yang, and Junchi Yan · 2022
Later among the works it cites.
Zeroth-order regularized optimization (zoro): Approximately sparse gradients and adaptive sampling
HanQin Cai, Daniel McKenzie, Wotao Yin, and Zhenliang Zhang · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Later among the works it cites.
Deepzero: Scaling up zeroth-order optimization for deep model training
Aochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer, Jiancheng Liu, Konstantinos Parasyris, Yihua Zhang, Zheng Zhang, Bhavya Kailkhura, and Sijia Liu · 2023
Later among the works it cites.
Fine-tuning happens in tiny subspaces: Exploring intrinsic task-specific subspaces of pre-trained language models
Zhong Zhang, Bang Liu, and Junming Shao · 2023
Later among the works it cites.
Zeroth-order optimization with orthogonal random directions
David Kozak, Cesare Molinari, Lorenzo Rosasco, Luis Tenorio, and Silvia Villa · 2023
Later among the works it cites.
Flora: Low-rank adapters are secretly gradient compressors
Yongchang Hao, Yanshuai Cao, and Lili Mou · 2024
Closest in time.
Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark
Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D Lee, Wotao Yin, Mingyi Hong, et al · 2024
Closest in time.
Grass: Compute efficient low-memory llm training with structured sparse gradients
Aashiq Muhamed, Oscar Li, David Woodruff, Mona Diab, and Virginia Smith · 2024
Closest in time.
Variance-reduced zeroth-order methods for fine-tuning language models
Tanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman, and Wooseok Ha · 2024
Closest in time.
Sparse mezo: Less parameters for better performance in zeroth-order llm fine-tuning
Yong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng, Cho-Jui Hsieh, and Yang You · 2024
Closest in time.
Addax: Memory-efficient fine-tuning of language models with a combination of forward-backward and forward-only passes
Zeman Li, Xinwei Zhang, and Meisam Razaviyayn · 2024
Closest in time.