Fetching the paper…
Reading the bibliography…
In the pursuit of explaining implicit regularization in deep learning, prominent focus was given to matrix and tensor factorizations, which correspond to simplified neural networks.
The expression of a tensor or a polyadic as a sum of products
Hitchcock, F. L · 1927
Earlier work this paper cites.
Numerical operator calculus in higher dimensions
Beylkin, G. and Mohlenkamp, M. J · 2002
Earlier work this paper cites.
Multiresolution quantum chemistry in multiwavelet bases
Harrison, R. J., Fann, G. I., Yanai, T., and Beylkin, G · 2003
Earlier work this paper cites.
On the efficient evaluation of coalescence integrals in population balance models
Hackbusch, W · 2006
Earlier work this paper cites.
Multilinear operators for higher-order decompositions
Kolda, T. G · 2006
Earlier work this paper cites.
Multivariate regression and machine learning with sums of separable functions
Beylkin, G., Garcke, J., and Mohlenkamp, M. J · 2009
Earlier work this paper cites.
A new scheme for the tensor representation
Hackbusch, W. and Kühn, S · 2009
Earlier work this paper cites.
Tensor decompositions and applications
Kolda, T. G. and Bader, B. W · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Hierarchical singular value decomposition of tensors
Grasedyck, L · 2010
Earlier work this paper cites.
Tensor spaces and numerical tensor calculus , volume 42
Hackbusch, W · 2012
Earlier work this paper cites.
Ordinary differential equations and dynamical systems , volume 140
Teschl, G · 2012
Earlier work this paper cites.
A literature survey of low-rank tensor approximation techniques
Grasedyck, L., Kressner, D., and Tobler, C · 2013
Earlier work this paper cites.
Simnets: A generalization of convolutional networks
Cohen, N. and Shashua, A · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Optimization on the hierarchical tucker manifold–applications to tensor completion
Da Silva, C. and Herrmann, F. J · 2015
Earlier work this paper cites.
Convolutional rectifier networks as generalized tensor decompositions
Cohen, N. and Shashua, A · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Tensorial mixture models
Sharir, O., Tamari, R., Cohen, N., and Shashua, A · 2016
Earlier work this paper cites.
Riemannian optimization for high-dimensional tensor completion
Steinlechner, M · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L · 2016
Earlier work this paper cites.
Inductive bias of deep convolutional networks through pooling geometry
Cohen, N. and Shashua, A · 2017
Earlier work this paper cites.
Analysis and design of convolutional networks via hierarchical tensor decompositions
Cohen, N., Sharir, O., Levine, Y., Tamari, R., Yakira, D., and Shashua, A · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
Neyshabur, B · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Low rank tensor recovery via iterative hard thresholding
Rauhut, H., Schneider, R., and Stojanac, Ž · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Earlier work this paper cites.
A tensor analysis on dense connectivity via convolutional arithmetic circuits
Balda, E. R., Behboodi, A., and Mathar, R · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Bartlett, P., Helmbold, D., and Long, P · 2018
Cited alongside, same era.
Boosting dilated convolutional networks with mixed tensor decompositions
Cohen, N., Tamari, R., and Shashua, A · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Du, S. S., Hu, W., and Lee, J. D · 2018
Cited alongside, same era.
Hierarchical quantum classifiers
Grant, E., Benedetti, M., Cao, S., Hallam, A., Lockhart, J., Stojevic, V., Green, A. G., and Severini, S · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Cited alongside, same era.
Expressive power of recurrent neural networks
Unique properties of wide minima in deep networks
Mulayoff, R. and Michaeli, T · 2020
Later among the works it cites.
Implicit regularization in deep learning may not be explainable by norms
Razin, N. and Cohen, N · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
On the implicit bias of initialization shape: Beyond infinitesimal mirror descent
Azulay, S., Moroshko, E., Nacson, M. S., Woodworth, B., Srebro, N., Globerson, A., and Soudry, D · 2021
Later among the works it cites.
More is less: Inducing sparsity via overparameterization
Chou, H.-H., Maly, J., and Rauhut, H · 2021
Later among the works it cites.
Limitations of implicit bias in matrix sensing: Initialization rank matters
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khrulkov, V., Novikov, A., and Oseledets, I · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Cited alongside, same era.
Learning long-range spatial dependencies with horizontal gated recurrent units
Linsley, D., Kim, J., Veerabadran, V., Windolf, C., and Serre, T · 2018
Cited alongside, same era.
On the expressive power of overlapping architectures of deep learning
Sharir, O. and Shashua, A · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Learning relevant features of data with multi-scale tensor networks
Stoudenmire, E. M · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Cited alongside, same era.
Eftekhari, A. and Zygalakis, K · 2021
Later among the works it cites.
Continuous vs. discrete optimization of deep neural networks
Elkabetz, O. and Cohen, N · 2021
Later among the works it cites.
Implicit convex regularizers of cnn architectures: Convex optimization of two-and three-layer networks in polynomial time
Ergen, T. and Pilanci, M · 2021
Later among the works it cites.
Quantum-inspired machine learning on high-energy physics data
Felser, T., Trenti, M., Sestini, L., Gianelle, A., Zuliani, D., Lucchesi, D., and Montangero, S · 2021
Later among the works it cites.
Understanding deflation process in over-parametrized tensor decomposition
Ge, R., Ren, Y., Wang, X., and Zhou, M · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z., Wei, C., Lee, J., and Ma, T · 2021
Later among the works it cites.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
Jagadeesan, M., Razenshteyn, I., and Gunasekar, S · 2021
Later among the works it cites.
Supervised learning and canonical decomposition of multivariate functions
Kargas, N. and Sidiropoulos, N. D · 2021
Later among the works it cites.
Geometry of linear convolutional networks
Kohn, K., Merkh, T., Montúfar, G., and Trager, M · 2021
Later among the works it cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Lyu, K., Li, Z., Wang, R., and Arora, S · 2021
Later among the works it cites.
Implicit regularization in deep tensor factorization
Milanesi, P., Kadri, H., Ayache, S., and Artières, T · 2021
Later among the works it cites.
On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks
Min, H., Tarmoun, S., Vidal, R., and Mallada, E · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Mulayoff, R., Michaeli, T., and Soudry, D · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Pesme, S., Pillaud-Vivien, L., and Flammarion, N · 2021
Later among the works it cites.
Implicit regularization in tensor factorization
Razin, N., Maman, A., and Cohen, N · 2021
Later among the works it cites.
Towards understanding learning in neural networks with linear teachers
Sarussi, R., Brutzkus, A., and Globerson, A · 2021
Later among the works it cites.
A theoretical analysis of fine-tuning with linear teachers
Shachaf, G., Brutzkus, A., and Globerson, A · 2021
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Later among the works it cites.
Implicit regularization in relu networks with the square loss
Vardi, G. and Shamir, O · 2021
Later among the works it cites.
On margin maximization in linear and relu networks
Vardi, G., Shamir, O., and Srebro, N · 2021
Later among the works it cites.
Which transformer architecture fits my data? a vocabulary bottleneck in self-attention
Wies, N., Levine, Y., Jannai, D., and Shashua, A · 2021
Later among the works it cites.
A unifying view on implicit bias in training linear neural networks
Yun, C., Krishnan, S., and Mobahi, H · 2021
Later among the works it cites.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bah, B., Rauhut, H., Terstiege, U., and Westdickenberg, M · 2022
Closest in time.
The inductive bias of in-context learning: Rethinking pretraining example design
Levine, Y., Wies, N., Jannai, D., Navon, D., Hoshen, Y., and Shashua, A · 2022
Closest in time.