Fetching the paper…
Reading the bibliography…
Despite the dominance and effectiveness of scaling, resulting in large networks with hundreds of billions of parameters, the necessity to train overparameterized models remains poorly understood, while training costs grow exponentially.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. J. Hsu, S. Ma, and S. Mandal · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Mixed precision training
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Earlier work this paper cites.
Implicit regularization in deep matrix factorization, 2019
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Earlier work this paper cites.
Stabilizing the lottery ticket hypothesis
J. Frankle, G. Karolina Dziugaite, D. M. Roy, and M. Carbin · 2019
Earlier work this paper cites.
Deep double descent: where bigger models and more data hurt
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Cited alongside, same era.
Low-rank bottleneck in multi-head attention models
S. Bhojanapalli, C. Yun, A. S. Rawat, S. Reddi, and S. Kumar · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Low-rank compression of neural nets: Learning the rank of each layer
Y. Idelbayev and M. A. Carreira-Perpinan · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. J. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. De Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. Rae, E. Elsen, and L. Sifre · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. M. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. C. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. García, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Díaz, O. Firat, M. Catasta, J. Wei, K. S. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with IO-awareness
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization through memorization: Nearest neighbor language models
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2020
Cited alongside, same era.
Hotcake: Higher order tucker articulated kernels for deeper cnn compression
R. Lin, C.-Y. Ko, Z. He, C. Chen, Y. Cheng, H. Yu, G. Chesi, and N. Wong · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
N. Shazeer · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
A. Aghajanyan, S. Gupta, and L. Zettlemoyer · 2021
Cited alongside, same era.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Re · 2022
Later among the works it cites.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Later among the works it cites.
Krona: Parameter efficient tuning with kronecker adapter
A. Edalati, M. S. Tahaei, I. Kobyzev, V. Nia, J. J. Clark, and M. Rezagholizadeh · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. W. Rae, and L. Sifre · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Later among the works it cites.
Exploring low rank training of deep neural networks
S. R. Kamalakara, A. F. Locatelli, B. Venkitesh, J. Ba, Y. Gal, and A. N. Gomez · 2022
Later among the works it cites.
Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations
S. Schotthöfer, E. Zangrando, J. Kusch, G. Ceruti, and F. Tudisco · 2022
Later among the works it cites.
Lst: Ladder side-tuning for parameter and memory efficient transfer learning
Y.-L. Sung, J. Cho, and M. Bansal · 2022
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Closest in time.
Scaling down to scale up: A guide to parameter-efficient fine-tuning, 2023
V. Lialin, V. Deshpande, and A. Rumshisky · 2023
Closest in time.
ELRT: Towards efficient low-rank training for compact neural networks, 2023
Y. Sui, M. Yin, W. Yang, Y. Gong, J. Xiao, H. Phan, D. Ding, X. Xu, S. Liu, Z. Chen, and B. Yuan · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
Inrank: Incremental low-rank learning
J. Zhao, Y. Zhang, B. Chen, F. Schäfer, and A. Anandkumar · 2023
Closest in time.