Fetching the paper…
Reading the bibliography…
We report a series of robust empirical observations, demonstrating that deep Neural Networks learn the examples in both the training and test sets in a similar order.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., Williams, R. J., et al · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M., and White, H · 1989
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 1989
Earlier work this paper cites.
Principal components, minor components, and linear neural networks
Oja, E · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al · 1998
Earlier work this paper cites.
Active learning for logistic regression: an evaluation
Schein, A. I. and Ungar, L. H · 2007
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Multi-class adaboost
Hastie, T., Rosset, S., Zhu, J., and Zou, H · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Recognizing indoor scenes
Quattoni, A. and Torralba, A · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Erhan, D., Bengio, Y., Courville, A., Manzagol, P.-A., Vincent, P., and Bengio, S · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Self-paced learning for latent variable models
Kumar, M. P., Packer, B., and Koller, D · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Tiny imagenet visual recognition challenge
Le, Y. and Yang, X · 2015
Cited alongside, same era.
Understanding image representations by measuring their equivariance and equivalence
Lenc, K. and Vedaldi, A · 2015
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. E · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Deep nets don’t learn via memorization
Krueger, D., Ballas, N., Jastrzebski, S., Arpit, D., Kanwal, M. S., Maharaj, T., Bengio, E., Fischer, A., and Courville, A · 2017
Later among the works it cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Later among the works it cites.
Vggface2: A dataset for recognising faces across pose and age
Cao, Q., Shen, L., Xie, W., Parkhi, O. M., and Zisserman, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding neural networks through deep visualization
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. E · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
Shrivastava, A., Gupta, A., and Girshick, R · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H., Heafield, K., Neckermann, T., Seide, F., Germann, U., Aji, A. F., Bogoychev, N., Martins, A. F. T., and Birch, A · 2018
Later among the works it cites.
Insights on representational similarity in neural networks with canonical correlation
Morcos, A., Raghu, M., and Bengio, S · 2018
Later among the works it cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2018
Later among the works it cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Wang, L., Hu, L., Gu, J., Hu, Z., Wu, Y., He, K., and Hopcroft, J · 2018
Later among the works it cites.
Separability and geometry of object manifolds in deep neural networks
Cohen, U., Chung, S., Lee, D. D., and Sompolinsky, H · 2019
Closest in time.
On the power of curriculum learning in training deep networks
Hacohen, G. and Weinshall, D · 2019
Closest in time.
Learning to combine grammatical error corrections
Kantor, Y., Katz, Y., Choshen, L., Cohen-Karlik, E., Liberman, N., Toledo, A., Menczel, A., and Slonim, N · 2019
Closest in time.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Closest in time.
Stack overflow bigquery program language dataset
Public domain dataset a · 2019
Closest in time.
Tiny imagenet challenge
Public domain dataset b · 2019
Closest in time.