Fetching the paper…
Reading the bibliography…
This work attempts to provide a plausible theoretical framework that aims to interpret modern deep (convolutional) networks from the principles of data compression and discriminative representation.
Neural collaborative subspace clustering
Tong Zhang, Pan Ji, Mehrtash Harandi, Wenbing Huang, and Hongdong Li · 1904
Earlier work this paper cites.
A rate-distortion framework for explaining neural network decisions
Jan MacDonald, Stephan Wäldchen, Sascha Hauch, and Gitta Kutyniok · 1905
Earlier work this paper cites.
A logical calculus of the ideas immanent in nervous activity
W. McCulloch and W. Pitts · 1943
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie · 1985
Earlier work this paper cites.
Induction of decision trees
J. R. Quinlan · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
A. Barron · 1991
Earlier work this paper cites.
Nonlinear principal component analysis using autoassociative neural networks
Mark A Kramer · 1991
Earlier work this paper cites.
The highly irregular firing of cortical cells is inconsistent with temporal integration of random EPSPs
William R Softky and Christof Koch · 1993
Earlier work this paper cites.
An introduction to the conjugate gradient method without the agonizing pain
Jonathan R Shewchuk · 1994
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun and Yoshua Bengio · 1995
Earlier work this paper cites.
Learning algorithms for classification: A comparison on handwritten digit recognition
Yann LeCun, Lawrence D Jackel, Léon Bottou, Corinna Cortes, John S Denker, Harris Drucker, Isabelle Guyon, Urs A Muller, Eduard Sackinger, Patrice Simard, et al · 1995
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Yann LeCun · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Independent component analysis: algorithms and applications
A. Hyvärinen and E. Oja · 2000
Earlier work this paper cites.
Principal Component Analysis
Ian T Jolliffe · 2002
Earlier work this paper cites.
Cluster ensembles—a knowledge reuse framework for combining multiple partitions
Alexander Strehl and Joydeep Ghosh · 2002
Earlier work this paper cites.
Neural Engineering: Computation, Representation and Dynamics in Neurobiological Systems
Chris Eliasmith and Charles Anderson · 2003
Earlier work this paper cites.
Log-det heuristic for matrix rank minimization with applications to Hankel and Euclidean distance matrices
M. Fazel, H. Hindi, and S. P. Boyd · 2003
Earlier work this paper cites.
Convex optimization
Stephen P Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Yann LeCun, Fu Jie Huang, and Leon Bottou · 2004
Earlier work this paper cites.
Intrinsic dimensionality estimation of submanifolds in rd
Matthias Hein and Jean-Yves Audibert · 2005
Earlier work this paper cites.
The multiscale structure of non-differentiable image manifolds
Michael B Wakin, David L Donoho, Hyeokho Choi, and Richard G Baraniuk · 2005
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Segmentation of multivariate mixed data via lossy data coding and compression
Yi Ma, Harm Derksen, Wei Hong, and John Wright · 2007
Earlier work this paper cites.
Low-frequency local field potentials and spikes in primary visual cortex convey independent visual information
Andrei Belitski, Arthur Gretton, Cesare Magri, Yusuke Murayama, Marcelo A. Montemurro, Nikos K. Logothetis, and Stefano Panzeri · 2008
Earlier work this paper cites.
Classification via minimum incremental coding length (MICL)
John Wright, Yangyu Tao, Zhouchen Lin, Yi Ma, and Heung-Yeung Shum · 2008
Earlier work this paper cites.
Robust face recognition via sparse representation
John Wright, Allen Y. Yang, Arvind Ganesh, S. Shankar Sastry, and Yi Ma · 2008
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Learning invariant features through topographic filter maps
Koray Kavukcuoglu, Marc’Aurelio Ranzato, Rob Fergus, and Yann LeCun · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Learning fast approximations of sparse coding
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
Evaluation of pooling operations in convolutional architectures for object recognition
Dominik Scherer, Andreas Müller, and Sven Behnke · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee · 2011
Earlier work this paper cites.
Contractive auto-encoders: Explicit invariance during feature extraction
Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio · 2011
Earlier work this paper cites.
On circulant matrices
Irwin Kra and Santiago R Simanca · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
The cosparse analysis model and algorithms
S. Nam, M.E. Davies, M. Elad, and R. Gribonval · 2012
Earlier work this paper cites.
Learning with recursive perceptual representations
Oriol Vinyals, Yangqing Jia, Li Deng, and Trevor Darrell · 2012
Earlier work this paper cites.
Toward a practical face recognition system: Robust alignment and illumination by sparse representation
Andrew Wagner, John Wright, Arvind Ganesh, Zihan Zhou, Hossein Mobahi, and Yi Ma · 2012
Earlier work this paper cites.
Invariant scattering convolution networks
Joan Bruna and Stéphane Mallat · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Fast training of convolutional networks through FFTs, 2013
Michael Mathieu, Mikael Henaff, and Yann LeCun · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Dictionary learning for analysis-synthesis thresholding
R. Rubinstein and M. Elad · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
PCANet: A simple deep learning baseline for image classification?
Tsung-Han Chan, Kui Jia, Shenghua Gao, Jiwen Lu, Zinan Zeng, and Yi Ma · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Logdet rank minimization with application to subspace clustering
Zhao Kang, Chong Peng, Jie Cheng, and Qiang Cheng · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Simple learned weighted sums of inferior temporal neuronal firing rates accurately predict human core object recognition performance
Najib J. Majaj, Ha Hong, Ethan A. Solomon, and James J. DiCarlo · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
Caveats for information bottleneck in deterministic scenarios
Artemy Kolchinsky, Brendan D Tracey, and Steven Van Kuyk · 2018
Later among the works it cites.
OLE: Orthogonal low-rank embedding-a plug and play geometric loss for deep learning
José Lezama, Qiang Qiu, Pablo Musé, and Guillermo Sapiro · 2018
Later among the works it cites.
Deep learning for universal linear embeddings of nonlinear dynamics
Bethany Lusch, J. Kutz, and Steven Brunton · 2018
Later among the works it cites.
On the implicit bias of dropout
Poorya Mianjy, Raman Arora, and Rene Vidal · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Fast convolutional nets with fbfft: A GPU performance evaluation
Nicolas Vasilache, J. Johnson, Michaël Mathieu, Soumith Chintala, Serkan Piantino, and Y. LeCun · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Cited alongside, same era.
Optimization Techniques in Computer Vision
Mongi A Abidi, Andrei V Gribok, and Joonki Paik · 2016
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Cited alongside, same era.
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Later among the works it cites.
Multilayer convolutional sparse modeling: Pursuit and dictionary learning
Jeremias Sulam, Vardan Papyan, Yaniv Romano, and Michael Elad · 2018
Later among the works it cites.
A mathematical theory of deep convolutional neural networks for feature extraction
T. Wiatowski and H. Bölcskei · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin · 2018
Later among the works it cites.
Scalable deep k-subspace clustering
Tong Zhang, Pan Ji, Mehrtash Harandi, Richard Hartley, and Ian Reid · 2018
Later among the works it cites.
Deep adversarial subspace clustering
Pan Zhou, Yunqing Hou, and Jiashi Feng · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Later among the works it cites.
A general theory of equivariant CNNs on homogeneous spaces
Taco S Cohen, Mario Geiger, and Maurice Weiler · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2019
Later among the works it cites.
Automatic Machine Learning: Methods, Systems, Challenges
Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren, editors · 2019
Later among the works it cites.
Invariant information clustering for unsupervised image classification and segmentation
Xu Ji, João F Henriques, and Andrea Vedaldi · 2019
Later among the works it cites.
Multichannel sparse blind deconvolution on the sphere
Yanjun Li and Yoram Bresler · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing
Vishal Monga, Yuelong Li, and Yonina C Eldar · 2019
Later among the works it cites.
A decoder-free approach for unsupervised clustering and manifold learning with random triplet mining
Oliver Nina, Jamison Moody, and Clarissa Milligan · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
A nonconvex approach for exact and efficient multichannel sparse blind deconvolution
Qing Qu, Xiao Li, and Zhihui Zhu · 2019
Later among the works it cites.
The singular values of convolutional layers
Hanie Sedghi, Vineet Gupta, and Philip M Long · 2019
Later among the works it cites.
Learning with bad training data via iterative trimmed loss minimization
Yanyao Shen and Sujay Sanghavi · 2019
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Later among the works it cites.
Deep comprehensive correlation mining for image clustering
Jianlong Wu, Keyu Long, Fei Wang, Chen Qian, Cheng Li, Zhouchen Lin, and Hongbin Zha · 2019
Later among the works it cites.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2020
Later among the works it cites.
Selecting the number of components in PCA via random signflips, 2020
David Hong, Yue Sheng, and Edgar Dobriban · 2020
Later among the works it cites.
Multimodal image synthesis with conditional implicit maximum likelihood estimation
Ke Li, Shichong Peng, Tianhao Zhang, and Jitendra Malik · 2020
Later among the works it cites.
Neural collapse with unconstrained features
Dustin G Mixon, Hans Parshall, and Jianzong Pi · 2020
Later among the works it cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Later among the works it cites.
Deep isometric learning for visual recognition
Haozhi Qi, Chong You, Xiaolong Wang, Yi Ma, and Jitendra Malik · 2020
Later among the works it cites.
Supervised deep sparse coding networks for image classification
Xiaoxia Sun, Nasser M Nasrabadi, and Trac D Tran · 2020
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Later among the works it cites.
On the optimal weighted ℓ 2 \ell_{2} regularization in overparameterized linear regression
Denny Wu and J. Xu · 2020
Later among the works it cites.
Rethinking bias-variance trade-off for generalization of neural networks
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma · 2020
Later among the works it cites.
Learning diverse and discriminative representations via the principle of maximal coding rate reduction
Yaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song, and Yi Ma · 2020
Later among the works it cites.
Deep network classification by scattering and homotopy dictionary learning
John Zarka, Louis Thiry, Tomás Angles, and Stéphane Mallat · 2020
Later among the works it cites.
A continual learning survey: Defying forgetting in classification tasks
Matthias Delange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Greg Slabaugh, and Tinne Tuytelaars · 2021
Closest in time.
Layer-peeled model: Toward understanding well-trained deep neural networks
Cong Fang, Hangfeng He, Qi Long, and Weijie J Su · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
A critique of self-expressive deep subspace clustering
Benjamin David Haeffele, Chong You, and Rene Vidal · 2021
Closest in time.
Neural collapse under mse loss: Proximity to and dynamics on the central path
XY Han, Vardan Papyan, and David L Donoho · 2021
Closest in time.
Convolutional normalization: Improving deep convolutional network robustness and training
Sheng Liu, Xiao Li, Yuexiang Zhai, Chong You, Zhihui Zhu, Carlos Fernandez-Granda, and Qing Qu · 2021
Closest in time.
The intrinsic dimension of images and its impact on learning
Phil Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum, and Tom Goldstein · 2021
Closest in time.
Are pre-trained convolutions better than pre-trained transformers?, 2021
Yi Tay, Mostafa Dehghani, Jai Gupta, Dara Bahri, Vamsi Aribandi, Zhen Qin, and Donald Metzler · 2021
Closest in time.
High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications
John Wright and Yi Ma · 2021
Closest in time.
Incremental learning via rate reduction
Ziyang Wu, Christina Baek, Chong You, and Yi Ma · 2021
Closest in time.
Separation and concentration in deep networks
John Zarka, Florentin Guth, and Stéphane Mallat · 2021
Closest in time.
A geometric analysis of neural collapse with unconstrained features, 2021
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Closest in time.