Fetching the paper…
Reading the bibliography…
Fine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadij Semenovič Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J.C. Spall · 1992
Earlier work this paper cites.
Optimal time and minimum space-time product for reversing a certain class of programs
José Grimm, Loīc Pottier, and Nicole Rostaing-Schmidt · 1996
Earlier work this paper cites.
A one-measurement form of simultaneous perturbation stochastic approximation
James C Spall · 1997
Earlier work this paper cites.
Building a question answering test collection
Ellen M Voorhees and Dawn M Tice · 2000
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan · 2007
Earlier work this paper cites.
A low rank approach to automatic differentiation
Hany S Abdel-Khalik, Paul D Hovland, Andrew Lyons, Tracy E Stover, and Jean Utke · 2008
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
Algorithm for stochastic approximation with trial input perturbation in the nonstationary problem of optimization
Alexander Timurovich Vakhitov, Oleg Nikolaevich Granichin, and LS Gurevich · 2009
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Maxim Raginsky and Alexander Rakhlin · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
Alekh Agarwal, Peter L. Bartlett, Pradeep Ravikumar, and Martin J. Wainwright · 2012
Earlier work this paper cites.
Query complexity of derivative-free optimization
Kevin G Jamieson, Robert Nowak, and Ben Recht · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Variance reduction for stochastic gradient optimization
Chong Wang, Xi Chen, Alexander J Smola, and Eric P Xing · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C. Duchi, Michael I. Jordan, Martin J. Wainwright, and Andre Wibisono · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh · 2017
Earlier work this paper cites.
On blackbox backpropagation and jacobian sensing
Krzysztof M Choromanski and Vikas Sindhwani · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir · 2017
Earlier work this paper cites.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Minimal effort back propagation for convolutional neural networks
Bingzhen Wei, Xu Sun, Xuancheng Ren, and Jingjing Xu · 2017
Cited alongside, same era.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates
Krishnakumar Balasubramanian and Saeed Ghadimi · 2018
Cited alongside, same era.
Adaptive sampling strategies for stochastic optimization
Raghu Bollapragada, Richard Byrd, and Jorge Nocedal · 2018
Cited alongside, same era.
Gradient-free multi-agent nonconvex nonsmooth optimization
Davood Hajinezhad and Michael M Zavlanos · 2018
Cited alongside, same era.
Dissecting hessian: Understanding common structure of hessian in neural networks
Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang, and Rong Ge · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney · 2020
Later among the works it cites.
Faster neural network training with approximate tensor operations
Menachem Adelman, Kfir Levy, Ido Hakimi, and Mark Silberstein · 2021
Later among the works it cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer · 2021
Later among the works it cites.
A zeroth-order block coordinate descent algorithm for huge-scale black-box optimization
HanQin Cai, Yuchen Lou, Daniel McKenzie, and Wotao Yin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Cited alongside, same era.
Zeroth-order stochastic variance reduction for nonconvex optimization
Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Paishun Ting, Shiyu Chang, and Lisa Amini · 2018
Cited alongside, same era.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Cited alongside, same era.
Stochastic zeroth-order optimization in high dimensions
Yining Wang, Simon Du, Sivaraman Balakrishnan, and Aarti Singh · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
Fairscale: A general purpose modular pytorch library for high performance and large scale training, 2021
FairScale authors · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Zhiyuan Li, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
A mathematical exploration of why language models help solve downstream tasks
Nikunj Saunshi, Sadhika Malladi, and Sanjeev Arora · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2021
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al · 2022
Later among the works it cites.
Zeroth-order nonconvex stochastic optimization: Handling constraints, high dimensionality, and saddle points
Krishnakumar Balasubramanian and Saeed Ghadimi · 2022
Later among the works it cites.
Gradients without backpropagation, 2022
Atılım Güneş Baydin, Barak A. Pearlmutter, Don Syme, Frank Wood, and Philip Torr · 2022
Later among the works it cites.
Zeroth-order regularized optimization (zoro): Approximately sparse gradients and adaptive sampling
HanQin Cai, Daniel McKenzie, Wotao Yin, and Zhenliang Zhang · 2022
Later among the works it cites.
Clip-tuning: Towards derivative-free prompt learning with a mixture of rewards
Yekun Chai, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
RLPrompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric Xing, and Zhiting Hu · 2022
Later among the works it cites.
Black-box prompt learning for pre-trained language models
Shizhe Diao, Xuechun Li, Yong Lin, Zhichao Huang, and Tong Zhang · 2022
Later among the works it cites.
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
Promptboosting: Black-box text classification with ten forward passes
Bairu Hou, Joe O’Connor, Jacob Andreas, Shiyu Chang, and Yang Zhang · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang · 2022
Later among the works it cites.
Robust training of neural networks using scale invariant architectures
Zhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank Reddi, and Sanjiv Kumar · 2022
Later among the works it cites.
What makes good in-context examples for GPT-3?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Later among the works it cites.
A kernel-based view of language model fine-tuning
Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen, and Sanjeev Arora · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Grips: Gradient-free, edit-based instruction search for prompting large language models
Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal · 2022
Later among the works it cites.
BBTv2: Towards a gradient-free future with large language models
Tianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou, Xuanjing Huang, and Xipeng Qiu · 2022
Later among the works it cites.
OpenAI · 2023
Closest in time.
Zeroth-order optimization meets human feedback: Provable learning via ranking oracles, 2023
Zhiwei Tang, Dmitry Rybin, and Tsung-Hui Chang · 2023
Closest in time.
Iterative forward tuning boosts in-context learning in language models, 2023
Jiaxi Yang, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li · 2023
Closest in time.