Fetching the paper…
Reading the bibliography…
We describe the emergence of a Convolution Bottleneck (CBN) structure in CNNs, where the network uses its first few layers to transform the input representation into a representation that is supported only along a few frequencies and channels, before using the last few layers to map back to the outputs.
Communication in the presence of noise
Shannon, C · 1949
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Kernel Methods for Deep Learning
Cho, Y. and Saul, L. K · 2009
Earlier work this paper cites.
Why are convolutional nets more sample-efficient than fully-connected nets?
Li, Z., Zhang, Y., and Arora, S · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Group invariant scattering
Mallat, S · 2012
Earlier work this paper cites.
Understanding deep neural networks with rectified linear units
Arora, R., Basu, A., Mianjy, P., and Mukherjee, A · 2016
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Universal approximations of invariant maps by neural networks
Yarotsky, D · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Cited alongside, same era.
Group invariance, stability to deformations, and complexity of deep convolutional representations
Bietti, A. and Mairal, J · 2019
Cited alongside, same era.
A general theory of equivariant cnns on homogeneous spaces
Cohen, T. S., Geiger, M., and Weiler, M · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Label noise sgd provably prefers flat global minimizers
Damian, A., Ma, T., and Lee, J. D · 2021
Later among the works it cites.
What happens after sgd reaches zero loss?–a mathematical framework
Li, Z., Wang, T., and Arora, S · 2021
Later among the works it cites.
Learning with invariances in random features and kernel models
Mei, S., Misiakiewicz, T., and Montanari, A · 2021
Later among the works it cites.
Relative stability toward diffeomorphisms indicates performance in deep nets
Petrini, L., Favero, A., Geiger, M., and Wyart, M · 2021
Later among the works it cites.
Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity, 2022
Jacot, A., Ged, F., Şimşek, B., Hongler, C., and Gabriel, F · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Cited alongside, same era.
Enhanced convolutional neural tangent kernels
Li, Z., Wang, R., Yu, D., Du, S. S., Hu, W., Salakhutdinov, R., and Arora, S · 2019
Cited alongside, same era.
Prevalence of neural collapse during the terminal phase of deep learning training
Papyan, V., Han, X., and Donoho, D. L · 2020
Cited alongside, same era.
Representation costs of linear neural networks: Analysis and design
Dai, Z., Karzand, M., and Srebro, N · 2021
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Gunasekar, S., Lee, J., Soudry, D., and Srebro, N
Cited in the paper.
Implicit bias of large depth networks: a notion of rank for nonlinear functions
Jacot, A
Cited in the paper.
Bottleneck structure in learned features: Low-dimension vs regularity tradeoff, 2023b
Jacot, A
Cited in the paper.
Understanding robustness and generalization of artificial neural networks through fourier masks
Karantzas, N., Besier, E., Ortega Caro, J., Pitkow, X., Tolias, A. S., Patel, A. B., and Anselmi, F · 2022
Later among the works it cites.
Learning with convolution and pooling operations in kernel methods
Misiakiewicz, T. and Mei, S · 2022
Later among the works it cites.
Synergy and symmetry in deep learning: Interactions between the data, model, and inference algorithm
Xiao, L. and Pennington, J · 2022
Later among the works it cites.
Theoretical analysis of the inductive biases in deep convolutional networks
Wang, Z. and Wu, L · 2023
Later among the works it cites.