Fetching the paper…
Reading the bibliography…
Despite large neural networks demonstrating remarkable abilities to complete different tasks, they require excessive memory usage to store the optimization states for training.
Approximate nearest neighbors: towards removing the curse of dimensionality
Indyk, P. and Motwani, R · 1998
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 1998
Earlier work this paper cites.
Experiments with random projection
Dasgupta, S · 2000
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Laurent, B. and Massart, P · 2000
Earlier work this paper cites.
Random projection in dimensionality reduction: applications to image and text data
Bingham, E. and Mannila, H · 2001
Earlier work this paper cites.
An elementary proof of a theorem of Johnson and Lindenstrauss
Dasgupta, S. and Gupta, A · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Nonnegative matrix factorization and I-divergence alternating minimization
Finesso, L. and Spreij, P · 2005
Earlier work this paper cites.
On variants of the Johnson–Lindenstrauss lemma
Matoušek, J · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Simple and deterministic matrix sketching
Liberty, E · 2013
Earlier work this paper cites.
Variance reduction for stochastic gradient optimization
Wang, C., Chen, X., Smola, A. J., and Xing, E. P · 2013
Earlier work this paper cites.
A tutorial on principal component analysis
Shlens, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Overview of the IWSLT 2017 evaluation campaign
Cettolo, M., Federico, M., Bentivogli, L., Niehues, J., Stüker, S., Sudoh, K., Yoshino, K., and Federmann, C · 2017
Earlier work this paper cites.
Why momentum really works
Goh, G · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Cited alongside, same era.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Cited alongside, same era.
Randomized automatic differentiation
Oktay, D., McGreivy, N., Aduol, J., Beatson, A., and Adams, R. P · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
FetchSGD: Communication-efficient federated learning with sketching
Rothchild, D., Panda, A., Ullah, E., Ivkin, N., Stoica, I., Braverman, V., Gonzalez, J., and Arora, R · 2020
Later among the works it cites.
8-bit optimizers via block-wise quantization
Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training Gopher
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Post, M · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V · 2018
Cited alongside, same era.
Momentum-based variance reduction in non-convex SGD
Cutkosky, A. and Orabona, F · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Later among the works it cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Later among the works it cites.
Towards understanding how momentum improves generalization in deep learning
Jelassi, S. and Li, Y · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Sketchy: Memory-efficient adaptive regularization with frequent directions
Feinberg, V., Chen, X., Sun, Y. J., Anil, R., and Hazan, E · 2023
Later among the works it cites.
Stack more layers differently: High-rank training through low-rank updates
Lialin, V., Shivagunde, N., Muckatira, S., and Rumshisky, A · 2023
Later among the works it cites.
Full parameter fine-tuning for large language models with limited resources
Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., and Qiu, X · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
GaLore: Memory-efficient LLM training by gradient low-rank projection
Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., and Tian, Y · 2024
Closest in time.