Fetching the paper…
Reading the bibliography…
The rank of neural networks measures information flowing across layers.
Das asymptotische verteilungsgesetz der eigenwerte linearer partieller differentialgleichungen
H. Weyl · 1912
Earlier work this paper cites.
Relations between two sets of variates
H. Hotelling · 1936
Earlier work this paper cites.
Products of random matrices
H. Furstenberg and H. Kesten · 1960
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
V. A. Marčenko and L. A. Pastur · 1967
Earlier work this paper cites.
Linear algebra
K. Hoffman · 1971
Earlier work this paper cites.
Subadditive ergodic theory
J. F. C. Kingman · 1973
Earlier work this paper cites.
Characteristic Lyapunov exponents and smooth ergodic theory
Y. B. Pesin · 1977
Earlier work this paper cites.
Random matrix products and measures on projective spaces
H. Furstenberg and H. Kesten · 1983
Earlier work this paper cites.
Determining Lyapunov exponents from a time series
A. Wolf, J. B. Swift, H. L. Swinney, and J. A. Vastano · 1985
Earlier work this paper cites.
The distribution of Lyapunov exponents: Exact results for random matrices
C. M. Newman · 1986
Earlier work this paper cites.
Principal component analysis
S. Wold, K. Esbensen, and P. Geladi · 1987
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Rank-r approximation of tensors using image-as-matrix representation
H. Wang and N. Ahuja · 2005
Earlier work this paper cites.
Generalized low rank approximations of matrices
J. Ye · 2005
Earlier work this paper cites.
Sparse principal component analysis
H. Zou, T. Hastie, and R. Tibshirani · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Numerical mathematics , volume 37
A. Quarteroni, R. Sacco, and F. Saleri · 2010
Earlier work this paper cites.
Differential topology , volume 33
M. W. Hirsch · 2012
Cited alongside, same era.
Low rank modeling of signed networks
C.-J. Hsieh, K.-Y. Chiang, and I. S. Dhillon · 2012
Cited alongside, same era.
Infinite-dimensional dynamical systems in mechanics and physics , volume 68
R. Temam · 2012
Cited alongside, same era.
On the equivalent of low-rank linear regressions and linear discriminant analysis based regressions
X. Cai, C. Ding, F. Nie, and H. Huang · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Lorslim: Low rank sparse linear methods for top-n recommendations
Y. Cheng, L. Yin, and Y. Yu · 2014
Cited alongside, same era.
Implicit regularization in deep matrix factorization
S. Arora, N. Cohen, W. Hu, and Y. Luo · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
From rank estimation to rank approximation: Rank residual constraint for image restoration
Z. Zha, X. Yuan, B. Wen, J. Zhou, J. Zhang, and C. Zhu · 2019
Later among the works it cites.
Is pruning compression?: Investigating pruning via network layer similarity
C. Blakeney, Y. Yan, and Z. Zong · 2020
Later among the works it cites.
Batch normalization provably avoids ranks collapse for randomly initialised deep networks
H. Daneshmand, J. Kohler, F. Bach, T. Hofmann, and A. Lucchi · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust low-rank tensor recovery: Models and algorithms
D. Goldfarb and Z. Qin · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Semi-supervised low-rank mapping learning for multi-label classification
L. Jing, L. Yang, J. Yu, and M. K. Ng · 2015
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Gaussian error linear units (GELUs)
D. Hendrycks and K. Gimpel · 2016
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Later among the works it cites.
Deep low-rank subspace clustering
M. Kheirandishfard, F. Zohrizadeh, and F. Kamangar · 2020
Later among the works it cites.
Hrank: Filter pruning using high-rank feature map
M. Lin, R. Ji, Y. Wang, Y. Zhang, B. Zhang, Y. Tian, and L. Shao · 2020
Later among the works it cites.
Graph-based non-convex low-rank regularization for image compression artifact reduction
J. Mu, R. Xiong, X. Fan, D. Liu, F. Wu, and W. Gao · 2020
Later among the works it cites.
GLU variants improve transformer
N. Shazeer · 2020
Later among the works it cites.
Practical low-rank communication compression in decentralized deep learning
T. Vogels, S. P. Karimireddy, and M. Jaggi · 2020
Later among the works it cites.
Low-rank constraints for fast inference in structured models
J. Chiu, Y. Deng, and A. Rush · 2021
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Y. Dong, J.-B. Cordonnier, and A. Loukas · 2021
Later among the works it cites.
Language model compression with weighted low-rank factorization
Y.-C. Hsu, T. Hua, S. Chang, Q. Lou, Y. Shen, and H. Jin · 2021
Later among the works it cites.
Compacter: Efficient low-rank hypercomplex adapter layers
R. Karimi Mahabadi, J. Henderson, and S. Ruder · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
C. H. Martin and M. W. Mahoney · 2021
Later among the works it cites.
ResMLP: Feedforward networks for image classification with data-efficient training
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, G. Izacard, A. Joulin, G. Synnaeve, J. Verbeek, et al · 2021
Later among the works it cites.