Fetching the paper…
Reading the bibliography…
Dense linear layers are the dominant computational bottleneck in foundation models.
Finitely correlated states on quantum spin chains
Fannes, M., Nachtergaele, B., and Werner, R. F · 1992
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 2002
Earlier work this paper cites.
Variational Learning of Inducing Variables in Sparse Gaussian Processes
Titsias, M. K · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Tensor-Train Decomposition
Oseledets, I. V · 2011
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., , and Hinton, G. E · 2012
Earlier work this paper cites.
Scalable inference for structured Gaussian process models
Saatçi, Y · 2012
Earlier work this paper cites.
Gaussian process kernels for pattern discovery and extrapolation
Wilson, A. and Adams, R · 2013
Earlier work this paper cites.
Algebraic geometry of matrix product states
Critch, A., Morton, J., et al · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma, J. B · 2015
Earlier work this paper cites.
Tensorizing Neural Networks
Novikov, A., Podoprikhin, D., Osokin, A., and Vetrov, D · 2015
Earlier work this paper cites.
Kernel interpolation for scalable structured gaussian processes (kiss-gp)
Wilson, A. and Nickisch, H · 2015
Earlier work this paper cites.
Layer Normalization
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N. and Welling, M · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Learning Efficient Convolutional Networks through Network Slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Cited alongside, same era.
FiLM: Visual Reasoning with a General Conditioning Layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks: w
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Pytorch image models
Wightman, R · 2019
Cited alongside, same era.
MLP-Mixer: An all-MLP Architecture for Vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., Lucic, M., and Dosovitskiy, A · 2021
Later among the works it cites.
Feature Learning in Infinite-Width Neural Networks
Yang, G. and Hu, E. J · 2021
Later among the works it cites.
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2021
Later among the works it cites.
Monarch: Expressive Structured Matrices for Efficient and Accurate Training
Dao, T., Chen, B., Sohoni, N., Desai, A., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Cited alongside, same era.
Generalizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data
Finzi, M., Stanton, S., Izmailov, P., and Wilson, A. G · 2020
Cited alongside, same era.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Cited alongside, same era.
Query-key normalization for transformers
Henry, A., Dachapally, P. R., Pawar, S., and Chen, Y · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
A unified weight initialization paradigm for tensorial convolutional neural networks
Pan, Y., Su, Z., Liu, A., Jingquan, W., Li, N., and Xu, Z · 2022
Later among the works it cites.
Scaling laws from the data manifold dimension
Sharma, U. and Kaplan, J · 2022
Later among the works it cites.
Scaling mlps: A tale of inductive bias
Bachmann, G., Anagnostidis, S., and Hofmann, T · 2023
Later among the works it cites.
Efficient GPT Model Pre-training using Tensor Train Matrix Representation
Chekalina, V., Novikov, G., Gusak, J., Oseledets, I., and Panchenko, A · 2023
Later among the works it cites.
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture
Fu, D. Y., Arora, S., Grogan, J., Johnson, I., Eyuboglu, S., Thomas, A. W., Spector, B., Poli, M., Rudra, A., and Ré, C · 2023
Later among the works it cites.
Differentiable learning of generalized structured matrices for efficient deep neural networks
Lee, C. and Kim, H.-S · 2023
Later among the works it cites.
Relora: High-rank training through low-rank updates
Lialin, V., Muckatira, S., Shivagunde, N., and Rumshisky, A · 2023
Later among the works it cites.
The quantization model of neural scaling
Michaud, E. J., Liu, Z., Girit, U., and Tegmark, M · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K., Alemi, A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., et al · 2023
Later among the works it cites.
Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit
Yang, G. and Littwin, E · 2023
Later among the works it cites.
Lora+: Efficient low rank adaptation of large models
Hayou, S., Ghosh, N., and Yu, B · 2024
Closest in time.
Cola: Exploiting compositional structure for automatic and efficient numerical linear algebra
Potapczynski, A., Finzi, M., Pleiss, G., and Wilson, A. G · 2024
Closest in time.
Galore: Memory-efficient llm training by gradient low-rank projection
Zhao, J., Zhang, Z., Chen, B., Wang, Z., Anandkumar, A., and Tian, Y · 2024
Closest in time.