Fetching the paper…
Reading the bibliography…
Modern practice for training classification deepnets involves a Terminal Phase of Training (TPT), which begins at the epoch where training error first vanishes; During TPT, the training error stays effectively zero while training loss is pushed towards zero.
\JournalTitle
RA Fisher, The use of multiple measurements in taxonomic problems · 1936
Earlier work this paper cites.
\JournalTitle
CE Shannon, Probability of error for optimal codes in a gaussian channel · 1959
Earlier work this paper cites.
TW Anderson, An introduction to multivariate statistical analysis, (Wiley New York), Technical report (1962)
1962
Earlier work this paper cites.
\JournalTitle
AR Webb, D Lowe, The optimised internal representation of multilayer classifier networks performs nonlinear discriminant analysis · 1990
Earlier work this paper cites.
A Dembo, O Zeitouni, Large deviations techniques and applications (2nd ed.) (1998)
1998
Earlier work this paper cites.
\JournalTitle
T Strohmer, RW Heath, Grassmannian frames with applications to coding and communication · 2003
Earlier work this paper cites.
(Springer), pp. 630–641 (2003)
K Fischer, B Gärtner, M Kutz, Fast smallest-enclosing-ball computation in high dimensions in European Symposium on Algorithms · 2003
Earlier work this paper cites.
(Springer Science & Business Media), (2006)
EL Lehmann, JP Romano, Testing statistical hypotheses · 2006
Earlier work this paper cites.
A Krizhevsky, G Hinton, Learning multiple layers of features from tiny images, (Citeseer), Technical report (2009)
2009
Earlier work this paper cites.
(Ieee), pp. 248–255 (2009)
J Deng, et al., Imagenet: A large-scale hierarchical image database in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on · 2009
Earlier work this paper cites.
\JournalTitle
Y LeCun, C Cortes, C Burges, MNIST handwritten digit database · 2010
Earlier work this paper cites.
\JournalTitle
S Mallat, Group invariant scattering · 2012
Earlier work this paper cites.
\JournalTitle
J Bruna, S Mallat, Invariant scattering convolution networks · 2013
Earlier work this paper cites.
K Simonyan, A Zisserman, Very deep convolutional networks for large-scale image recognition (2014)
2014
Earlier work this paper cites.
H Monajemi, DL Donoho, ClusterJob: An automated system for painless and reproducible massive computational experiments ( https://github.com/monajemi/clusterjob ) (2015)
2015
Earlier work this paper cites.
(IEEE), pp. 1212–1216 (2015)
T Wiatowski, H Bölcskei, Deep convolutional neural networks based on semi-discrete frames in 2015 IEEE International Symposium on Information Theory (ISIT) · 2015
Earlier work this paper cites.
pp. 770–778 (2016)
K He, X Zhang, S Ren, J Sun, Deep residual learning for image recognition in Proceedings of the IEEE conference on computer vision and pattern recognition · 2016
Earlier work this paper cites.
(IEEE), pp. 2368–2373 (2016)
H Monajemi, DL Donoho, V Stodden, Making massive computational experiments painless in 2016 IEEE International Conference on Big Data (Big Data) · 2016
Cited alongside, same era.
pp. 2574–2582 (2016)
SM Moosavi-Dezfooli, A Fawzi, P Frossard, Deepfool: a simple and accurate method to fool deep neural networks in Proceedings of the IEEE conference on computer vision and pattern recognition · 2016
Cited alongside, same era.
pp. 2149–2158 (2016)
T Wiatowski, M Tschannen, A Stanic, P Grohs, H Bölcskei, Discrete deep feature extraction: A theory and new architectures in International Conference on Machine Learning · 2016
Cited alongside, same era.
L Sagun, L Bottou, Y LeCun, Eigenvalues of the Hessian in deep learning: Singularity and beyond (2016)
2016
Cited alongside, same era.
H Xiao, K Rasul, R Vollgraf, Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms (2017)
2017
Cited alongside, same era.
pp. 4700–4708 (2017)
pp. 1611–1619 (2019)
M Belkin, A Rakhlin, AB Tsybakov, Does data interpolation contradict statistical optimality? in The 22nd International Conference on Artificial Intelligence and Statistics · 2019
Later among the works it cites.
\JournalTitle
M Belkin, D Hsu, S Ma, S Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off · 2019
Later among the works it cites.
M Belkin, Beyond empirical risk minimization: the lessons of deep learning (2019) MIT CBMM colloquium recording
2019
Later among the works it cites.
pp. 8026–8037 (2019)
A Paszke, et al., Pytorch: An imperative style, high-performance deep learning library in Advances in neural information processing systems · 2019
Later among the works it cites.
\JournalTitle
H Monajemi, et al., Ambitious data science can be painless · 2019
Later among the works it cites.
(IEEE), pp. 1–8 (2019)
NM Müller, K Markert, Identifying mislabeled instances in classification datasets in 2019 International Joint Conference on Neural Networks (IJCNN) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G Huang, Z Liu, L Van Der Maaten, KQ Weinberger, Densely connected convolutional networks in Proceedings of the IEEE conference on computer vision and pattern recognition · 2017
Cited alongside, same era.
G Pleiss, et al., Memory-efficient implementation of DenseNets (2017)
2017
Cited alongside, same era.
(IEEE), pp. 2420–2424 (2017)
R Ekambaram, DB Goldgof, LO Hall, Finding label noise examples in large scale datasets in 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2017
Cited alongside, same era.
\JournalTitle
T Wiatowski, H Bölcskei, A mathematical theory of deep convolutional neural networks for feature extraction · 2017
Cited alongside, same era.
L Sagun, U Evci, VU Guney, Y Dauphin, L Bottou, Empirical analysis of the Hessian of over-parametrized neural networks (2017)
2017
Cited alongside, same era.
\JournalTitle
V Papyan, Y Romano, M Elad, Convolutional neural networks analyzed via convolutional sparse coding · 2017
Cited alongside, same era.
(JMLR. org), pp. 854–863 (2017)
M Cisse, P Bojanowski, E Grave, Y Dauphin, N Usunier, Parseval networks: Improving robustness to adversarial examples in Proceedings of the 34th International Conference on Machine Learning-Volume 70 · 2017
Cited alongside, same era.
Later among the works it cites.
V Papyan, Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet Hessianss (2019)
2019
Later among the works it cites.
B Ghorbani, S Krishnan, Y Xiao, An investigation into neural net optimization via Hessians eigenvalue density (2019)
2019
Later among the works it cites.
\JournalTitle
Y Romano, A Aberdam, J Sulam, M Elad, Adversarial noise attacks of deep learning architectures: Stability analysis via sparse-modeled signals · 2019
Later among the works it cites.
\JournalTitle
A Aberdam, J Sulam, M Elad, Multi-layer sparse coding: The holistic way · 2019
Later among the works it cites.
Directory of AI benchmarks ( https://benchmarks.ai/ ) (2020)
2020
Closest in time.
Stanford DAWN ( https://dawn.cs.stanford.edu/ ) (2020)
2020
Closest in time.
Who is best at X? ( https://rodrigob.github.io/are_we_there_yet/build/ ) (2020)
2020
Closest in time.
Forward and backward function hooks — PyTorch documentation ( https://pytorch.org/tutorials/beginner/former_torchies/nnft_tutorial.html#forward-and-backward-function-hooks ) (2020) [Online; accessed 21-June-2020]
2020
Closest in time.
\JournalTitle
J Sulam, A Aberdam, A Beck, M Elad, On multi-layer basis pursuit, efficient algorithms and convolutional neural networks · 2020
Closest in time.
A Aberdam, D Simon, M Elad, When and how can deep generative models be inverted? (2020)
2020
Closest in time.
O Deniz, A Pedraza, N Vallez, J Salido, G Bueno, Robustness to adversarial examples can be improved with overfitting (2020)
2020
Closest in time.