Fetching the paper…
Reading the bibliography…
Training deep neural networks in low rank, i.e.
The lottery ticket hypothesis at scale
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 1903
Earlier work this paper cites.
The difficulty of training sparse neural networks
Evci, U., Pedregosa, F., Gomez, A. N., and Elsen, E · 1906
Earlier work this paper cites.
A signal propagation perspective for pruning neural networks at initialization
Lee, N., Ajanthan, T., Gould, S., and Torr, P. H. S · 1906
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 1912
Earlier work this paper cites.
Rank, trace-norm and max-norm
Srebro, N. and Shraibman, A · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., and Koehn, P · 2013
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions, 2014
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift, 2015
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Layer normalization, 2016
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Convolutional neural networks with low-rank regularization, 2016
Tai, C., Xiao, T., Zhang, Y., Wang, X., and E, W · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning, 2018
Martin, C. H. and Mahoney, M. W · 2018
Later among the works it cites.
Implicit regularization in deep matrix factorization, 2019
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achille, A., Rovere, M., and Soatto, S · 2017
Cited alongside, same era.
On compressing deep models by low rank and sparse decomposition
Yu, X., Liu, T., Wang, X., and Tao, D · 2017
Cited alongside, same era.
Later among the works it cites.
Initialization and regularization of factorized neural layers
Khodak, M., Tenenholtz, N. A., Mackey, L., and Fusi, N · 2021
Later among the works it cites.
Pufferfish: Communication-efficient models at no extra cost, 2021
Wang, H., Agarwal, S., and Papailiopoulos, D · 2021
Later among the works it cites.