Fetching the paper…
Reading the bibliography…
Fine-tuning language models (LMs) has demonstrated success in a wide array of downstream tasks.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Improved zeroth-order variance reduced algorithms and analysis for nonconvex optimization
Ji, K., Wang, Z., Zhou, Y., and Liang, Y · 1910
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Raginsky, M. and Rakhlin, A · 2011
Earlier work this paper cites.
Query complexity of derivative-free optimization
Jamieson, K. G., Nowak, R., and Recht, B · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
Duchi, J., Jordan, M., Wainwright, M., and Wibisono, A · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Cited alongside, same era.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Sun, X., Ren, X., Ma, S., and Wang, H · 2017
Cited alongside, same era.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, J., and Sohl-Dickstein, J · 2017
Cited alongside, same era.
Minimal effort back propagation for convolutional neural networks
Wei, B., Sun, X., Ren, X., and Xu, J · 2017
Cited alongside, same era.
Zeroth-order stochastic variance reduction for nonconvex optimization
Liu, S., Kailkhura, B., Chen, P.-Y., Ting, P., Chang, S., and Amini, L · 2018
Cited alongside, same era.
Linear convergence of cyclic saga
Park, Y. and Ryu, E. K · 2020
Later among the works it cites.
Structured policy iteration for linear quadratic regulator
Park, Y., Rossi, R., Wen, Z., Wu, G., and Zhao, H · 2020
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2020
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Variance reduced training with stratified sampling for forecasting models
Lu, Y., Park, Y., Chen, L., Wang, Y., De Sa, C., and Foster, D · 2021
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Release strategies and the social impacts of language models, 2019
Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGuffie, K., and Wang, J · 2019
Cited alongside, same era.
SuperGLUE: a stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., and Subbiah, M · 2020
Cited alongside, same era.
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Fine-tuning large language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark, 2024
Zhang, Y., Li, P., Hong, J., Li, J., Zhang, Y., Zheng, W., Chen, P.-Y., Lee, J. D., Yin, W., Hong, M., Wang, Z., Liu, S., and Chen, T · 2024
Closest in time.