Fetching the paper…
Reading the bibliography…
Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers.
Distribution density, tails, and outliers in machine learning: Metrics and applications
Carlini, N., Erlingsson, U., and Papernot, N. (2019) · 1910
Earlier work this paper cites.
What do compressed deep neural networks forget?
Hooker, S., Courville, A., Clark, G., Dauphin, Y., and Frome, A. (2019) · 1911
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Elman, J. L. (1993) · 1993
Earlier work this paper cites.
Neural network learning control of robot manipulators using gradually increasing task difficulty
Sanger, T. D. (1994) · 1994
Earlier work this paper cites.
Predicting neural network accuracy from weights
Unterthiner, T., Keysers, D., Gelly, S., Bousquet, O., and Tolstikhin, I. (2020) · 2002
Earlier work this paper cites.
Zielinski, P., Krishnan, S., and Chatterjee, S. (2020) · 2003
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Earlier work this paper cites.
Characterising bias in compressed models
Hooker, S., Moorosi, N., Clark, G., Bengio, S., and Denton, E. (2020) · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2015) · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Eigenvalues of the Hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y. (2016) · 2016
Earlier work this paper cites.
BranchyNet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S., McDanel, B., and Kung, H.-T. (2016) · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al. (2017) · 2017
Earlier work this paper cites.
Generalization in deep learning
Kawaguchi, K., Kaelbling, L. P., and Bengio, Y. (2017) · 2017
Earlier work this paper cites.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y. (2017) · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Nocedal, J., Tang, P. T. P., Mudigere, D., and Smelyanskiy, M. (2017) · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2017) · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Pennington, J. and Bahri, Y. (2017) · 2017
Cited alongside, same era.
SVCCA: singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. (2017) · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate Bayesian inference
Stephan, M., Hoffman, M. D., Blei, D. M., et al. (2017) · 2017
Cited alongside, same era.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R. (2017) · 2017
Cited alongside, same era.
DNN or k-NN: That is the generalize vs. memorize question
Cohen, G., Sapiro, G., and Giryes, R. (2018) · 2018
Cited alongside, same era.
Universal transformers
Generalization bounds for deep convolutional neural networks
Long, P. M. and Sedghi, H. (2019) · 2019
Later among the works it cites.
Do deep neural networks learn shallow learnable examples first?
Mangalam, K. and Prabhu, V. (2019) · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. (2019) · 2019
Later among the works it cites.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J. (2019) · 2019
Later among the works it cites.
BatchEnsemble: An alternative approach to efficient ensemble and lifelong learning
Wen, Y., Tran, D., and Ba, J. (2019) · 2019
Later among the works it cites.
Estimating example difficulty using variance of gradients
Agarwal, C. and Hooker, S. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L. (2018) · 2018
Cited alongside, same era.
Multi-scale dense networks for resource efficient image classification
Huang, G., Chen, D., Li, T., Wu, F., van der Maaten, L., and Weinberger, K. (2018) · 2018
Cited alongside, same era.
Predicting the generalization gap in deep networks with margin distributions
Jiang, Y., Krishnan, D., Mobahi, H., and Bengio, S. (2018) · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R. (2018) · 2018
Cited alongside, same era.
Understanding deep learning performance through an examination of test set difficulty: A psychometric case study
Lalor, J. P., Wu, H., Munkhdalai, T., and Yu, H. (2018) · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2018) · 2018
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
Morcos, A. S., Raghu, M., and Bengio, S. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Deep k-NN for noisy labels
Bahri, D., Jiang, H., and Gupta, M. (2020) · 2020
Later among the works it cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C. (2020) · 2020
Later among the works it cites.
Let’s agree to agree: Neural networks share classification order on real datasets
Hacohen, G., Choshen, L., and Weinshall, D. (2020) · 2020
Later among the works it cites.
The surprising simplicity of the early-time learning dynamics of neural networks
Hu, W., Xiao, L., Adlam, B., and Pennington, J. (2020) · 2020
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S. (2020) · 2020
Later among the works it cites.
Big transfer (BiT): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N. (2020) · 2020
Later among the works it cites.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A. (2020) · 2020
Later among the works it cites.
Hyperparameter ensembles for robustness and uncertainty quantification
Wenzel, F., Snoek, J., Tran, D., and Jenatton, R. (2020) · 2020
Later among the works it cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Xin, J., Tang, R., Lee, J., Yu, Y., and Lin, J. (2020) · 2020
Later among the works it cites.
PyHessian: Neural networks through the lens of the Hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. (2020) · 2020
Later among the works it cites.
Characterizing structural regularities of labeled data in overparameterized models
Jiang, Z., Zhang, C., Talwar, K., and Mozer, M. C. (2021) · 2021
Closest in time.
Understanding the failure modes of out-of-distribution generalization
Nagarajan, V., Andreassen, A., and Neyshabur, B. (2021) · 2021
Closest in time.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S. (2021) · 2021
Closest in time.
On the geometry of generalization and memorization in deep neural networks
Stephenson, C., suchismita padhy, Ganesh, A., Hui, Y., Tang, H., and Chung, S. (2021) · 2021
Closest in time.
When do curricula work?
Wu, X., Dyer, E., and Neyshabur, B. (2021) · 2021
Closest in time.