Fetching the paper…
Reading the bibliography…
Convolutional Neural Networks (CNNs) dominate various computer vision tasks since Alex Krizhevsky showed that they can be trained effectively and reduced the top-5 error from 26.2 % to 15.3 % on the ImageNet large scale visual recognition challenge.
C. Dugas, Y. Bengio et al. , “Incorporating second-order functional knowledge for better option pricing,” in Advances in Neural Information Processing Systems 13 (NIPS) , T. K. Leen, T. G. Dietterich, and V. Tresp, Eds. MIT Press, 2001, pp. 472–478. [Online]. Available: http://papers.nips.cc/paper/1920-incorporating-second-order-functional-knowledge-for-better-option-pricing.pdf
1920
Earlier work this paper cites.
S. E. Fahlman and C. Lebiere, “The cascade-correlation learning architecture,” 1989. [Online]. Available: http://repository.cmu.edu/compsci/1938/
1938
Earlier work this paper cites.
W. S. McCulloch and W. Pitts, “A logical calculus of the ideas immanent in nervous activity,” The bulletin of mathematical biophysics , vol. 5, no. 4, pp. 115–133, 1943
1943
Earlier work this paper cites.
M. R. Garey, D. S. Johnson, and L. Stockmeyer, “Some simplified NP-complete graph problems,” Theoretical computer science , vol. 1, no. 3, pp. 237–267, 1976
1976
Earlier work this paper cites.
Y. Nesterov, “A method of solving a convex programming problem with convergence rate o (1/k2),” in Soviet Mathematics Doklady , vol. 27, no. 2, 1983, pp. 372–376
1983
Earlier work this paper cites.
P. J. M. van Laarhoven and E. H. L. Aarts, Simulated annealing . Dordrecht: Springer Netherlands, 1987, pp. 7–15. [Online]. Available: http://dx.doi.org/10.1007/978-94-015-7744-1_2
1987
Earlier work this paper cites.
S. E. Fahlman, “An empirical study of learning speed in back-propagation networks,” 1988. [Online]. Available: http://repository.cmu.edu/cgi/viewcontent.cgi?article=2799&context=compsci
1988
Earlier work this paper cites.
S. J. Hanson, “Meiosis networks.” in NIPS , 1989, pp. 533–541. [Online]. Available: http://papers.nips.cc/paper/227-meiosis-networks.pdf
1989
Earlier work this paper cites.
Y. LeCun, J. S. Denker et al. , “Optimal brain damage.” in NIPs , vol. 2, 1989, pp. 598–605. [Online]. Available: http://yann.lecun.com/exdb/publis/pdf/lecun-90b.pdf
1989
Earlier work this paper cites.
A. Waibel, T. Hanazawa et al. , “Phoneme recognition using time-delay neural networks,” IEEE transactions on acoustics, speech, and signal processing , vol. 37, no. 3, pp. 328–339, Aug. 1989. [Online]. Available: http://ieeexplore.ieee.org/document/21701/
1989
Earlier work this paper cites.
C. Charalambous, “Conjugate gradient algorithm for efficient training of artificial neural networks,” IEEE Proceedings G-Circuits, Devices and Systems , vol. 139, no. 3, pp. 301–310, 1992. [Online]. Available: http://ieeexplore.ieee.org/document/143326/
1992
Earlier work this paper cites.
S. J. Nowlan and G. E. Hinton, “Simplifying neural networks by soft weight-sharing,” Neural computation , vol. 4, no. 4, pp. 473–493, 1992. [Online]. Available: https://www.cs.toronto.edu/˜hinton/absps/sunspots.pdf
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3-4, pp. 229–256, 1992
1992
Earlier work this paper cites.
B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” in International Conference on Neural Networks . IEEE, 1993, pp. 293–299. [Online]. Available: http://ee.caltech.edu/Babak/pubs/conferences/00298572.pdf
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE transactions on neural networks , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
M. Ester, H.-P. Kriegel et al. , “A density-based algorithm for discovering clusters in large spatial databases with noise.” in Kdd , vol. 96, no. 34, 1996, pp. 226–231
1996
Earlier work this paper cites.
Y. LeCun, L. Bottou et al. , “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, Nov. 1998. [Online]. Available: http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf
1998
Earlier work this paper cites.
Y. A. LeCun, L. Bottou et al. , Efficient BackProp , ser. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer Berlin Heidelberg, 1998, vol. 1524, pp. 9–50. [Online]. Available: http://dx.doi.org/10.1007/3-540-49430-8
1998
Earlier work this paper cites.
L. Prechelt, Early Stopping - But When? Berlin, Heidelberg: Springer Berlin Heidelberg, 1998, pp. 55–69. [Online]. Available: http://dx.doi.org/10.1007/3-540-49430-8_3
1998
Earlier work this paper cites.
C. J. B. Yann LeCun, Corinna Cortes, “The MNIST database of handwritten digits,” 1998. [Online]. Available: http://yann.lecun.com/exdb/mnist/
1998
Earlier work this paper cites.
M. Ankerst, M. M. Breunig et al. , “OPTICS: Ordering points to identify the clustering structure,” in ACM Sigmod record , vol. 28, no. 2. ACM, 1999, pp. 49–60
1999
Earlier work this paper cites.
W. Duch and N. Jankowski, “Survey of neural transfer functions,” Neural Computing Surveys , vol. 2, no. 1, pp. 163–212, 1999. [Online]. Available: ftp://ftp.icsi.berkeley.edu/pub/ai/jagota/vol2_6.pdf
1999
Earlier work this paper cites.
“The training performed by qnstrn,” Aug. 2000. [Online]. Available: http://www1.icsi.berkeley.edu/Speech/faq/nn-train.html
2000
Earlier work this paper cites.
M. R. Garey and D. S. Johnson, Computers and intractability . wh freeman New York, 2002, vol. 29
2002
Earlier work this paper cites.
V. Kurkova and M. Sanguineti, “Comparison of worst case errors in linear and neural network approximation,” IEEE Transactions on Information Theory , vol. 48, no. 1, pp. 264–275, Jan. 2002. [Online]. Available: http://ieeexplore.ieee.org/abstract/document/971754/
2002
Earlier work this paper cites.
R. T. Ng and J. Han, “CLARANS: A method for clustering objects for spatial data mining,” IEEE transactions on knowledge and data engineering , vol. 14, no. 5, pp. 1003–1016, 2002
2002
Earlier work this paper cites.
K. O. Stanley and R. Miikkulainen, “Evolving neural networks through augmenting topologies,” Evolutionary computation , vol. 10, no. 2, pp. 99–127, 2002. [Online]. Available: http://www.mitpressjournals.org/doi/abs/10.1162/106365602320169811
2002
Earlier work this paper cites.
A. E. Eiben and J. E. Smith, Introduction to evolutionary computing . Springer, 2003, vol. 53. [Online]. Available: https://dx.doi.org/10.1007/978-3-662-44874-8
2003
Earlier work this paper cites.
R. F. Fei-Fei and P. Perona, “Caltech 101,” 2003. [Online]. Available: http://www.vision.caltech.edu/Image_Datasets/Caltech101/Caltech101.html
2003
Earlier work this paper cites.
L. Fei-Fei, R. Fergus, and P. Perona, “One-shot learning of object categories,” IEEE transactions on pattern analysis and machine intelligence , vol. 28, no. 4, pp. 594–611, Apr. 2006. [Online]. Available: http://vision.stanford.edu/documents/Fei-FeiFergusPerona2006.pdf
2006
Earlier work this paper cites.
A. P. Griffin, G. Holub, “Caltech 256,” 2006. [Online]. Available: http://www.vision.caltech.edu/Image_Datasets/Caltech256/
2006
Earlier work this paper cites.
J. Elson, J. J. Douceur et al. , “Asirra: A CAPTCHA that exploits interest-aligned manual image categorization,” in ACM Conference on Computer and Communications Security (CCS) , no. 14. Association for Computing Machinery, Inc., Oct. 2007. [Online]. Available: https://www.microsoft.com/en-us/research/publication/asirra-a-captcha-that-exploits-interest-aligned-manual-image-categorization/
2007
Earlier work this paper cites.
P. P. Greg Griffin, Alex Holub, “Caltech-256 object category dataset,” Apr. 2007. [Online]. Available: http://authors.library.caltech.edu/7694/
2007
Earlier work this paper cites.
M. Marszalek and C. Schmid, “Accurate object localization with shape masks,” in Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2007, pp. 1–8. [Online]. Available: http://ieeexplore.ieee.org/document/4270110/
2007
Earlier work this paper cites.
P. Golle, “Machine learning attacks against the Asirra CAPTCHA,” in ACM conference on Computer and communications security (CCS) , no. 15. ACM, 2008, pp. 535–542
2008
Earlier work this paper cites.
M. Marszałek, “INRIA annotations for Graz-02 (IG02),” Oct. 2008. [Online]. Available: http://lear.inrialpes.fr/people/marszalek/data/ig02/
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
J. Bergstra, G. Desjardins et al. , “Quadratic polynomials learn better image features,” Département d’Informatique et de Recherche Opérationnelle, Université de Montréal, Tech. Rep. 1337, 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Apr. 2009. [Online]. Available: https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf
2009
Earlier work this paper cites.
L. Kaufman and P. J. Rousseeuw, Finding groups in data: an introduction to cluster analysis . John Wiley & Sons, 2009, vol. 344
2009
Earlier work this paper cites.
K. O. Stanley, D. B. D’Ambrosio, and J. Gauci, “A hypercube-based encoding for evolving large-scale neural networks,” Artificial life , vol. 15, no. 2, pp. 185–212, 2009. [Online]. Available: http://ieeexplore.ieee.org/document/6792316/
2009
Earlier work this paper cites.
R. Behmo, P. Marcombes et al. , “Towards optimal naive Bayes nearest neighbor,” in European Conference on Computer Vision (ECCV) . Springer, 2010, pp. 171–184
2010
Earlier work this paper cites.
Y.-L. Boureau, J. Ponce, and Y. LeCun, “A theoretical analysis of feature pooling in visual recognition,” in International Conference on Machine Learning (ICML) , no. 27, 2010, pp. 111–118. [Online]. Available: http://yann.lecun.com/exdb/publis/pdf/boureau-icml-10.pdf
2010
Earlier work this paper cites.
A. Coates, H. Lee, and A. Y. Ng, “An analysis of single-layer networks in unsupervised feature learning,” Ann Arbor , vol. 1001, no. 48109, p. 2, 2010. [Online]. Available: http://cs.stanford.edu/˜acoates/papers/coatesleeng_aistats_2011.pdf
2010
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick et al. , “Object detection with discriminatively trained part-based models,” IEEE transactions on pattern analysis and machine intelligence , vol. 32, no. 9, pp. 1627–1645, 2010
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks.” in Aistats , vol. 9, 2010, pp. 249–256. [Online]. Available: http://jmlr.org/proceedings/papers/v9/glorot10a/glorot10a.pdf
2010
Earlier work this paper cites.
K. Kavukcuoglu, P. Sermanet et al. , “Learning convolutional feature hierarchies for visual recognition,” in Advances in Neural Information Processing Systems 23 (NIPS) , J. D. Lafferty, C. K. I. Williams et al. , Eds. Curran Associates, Inc., 2010, pp. 1090–1098. [Online]. Available: http://papers.nips.cc/paper/4133-learning-convolutional-feature-hierarchies-for-visual-recognition.pdf
2010
Earlier work this paper cites.
S. Risi, J. Lehman, and K. O. Stanley, “Evolving the placement and density of neurons in the hyperneat substrate,” in Conference on Genetic and evolutionary computation , no. 12. ACM, 2010, pp. 563–570
2010
Earlier work this paper cites.
A. Coates, H. Lee, and A. Y. Ng, “STL-10 dataset,” 2011. [Online]. Available: http://cs.stanford.edu/˜acoates/stl10
2011
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research , vol. 12, no. Jul, pp. 2121–2159, 2011. [Online]. Available: http://www.jmlr.org/papers/volume12/duchi11a/duchi11a.pdf
2011
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks.” in Aistats , vol. 15, no. 106, 2011, p. 275. [Online]. Available: http://www.jmlr.org/proceedings/papers/v15/glorot11a/glorot11a.pdf
2011
Earlier work this paper cites.
J. Han, J. Pei, and M. Kamber, Data mining: concepts and techniques . Elsevier, 2011
2011
Earlier work this paper cites.
A. Karpathy, “Lessons learned from manually classifying CIFAR-10,” Apr. 2011. [Online]. Available: http://karpathy.github.io/2011/04/27/manually-classifying-cifar10/
2011
Earlier work this paper cites.
Y. Netzer, T. Wang et al. , “Reading digits in natural images with unsupervised feature learning,” in NIPS workshop on deep learning and unsupervised feature learning , vol. 2011, no. 2, 2011, p. 5. [Online]. Available: http://ufldl.stanford.edu/housenumbers/nips2011_housenumbers.pdf
2011
Earlier work this paper cites.
Y. Netzer, T. Wang et al. , “The street view house numbers (SVHN) dataset,” 2011. [Online]. Available: http://ufldl.stanford.edu/housenumbers/
2011
Earlier work this paper cites.
P. Sermanet and Y. LeCun, “Traffic sign recognition with multi-scale convolutional networks,” in International Joint Conference on Neural Networks (IJCNN) , Jul. 2011, pp. 2809–2813. [Online]. Available: http://ieeexplore.ieee.org/document/6033589/
2011
Earlier work this paper cites.
2011
Earlier work this paper cites.
J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” Journal of Machine Learning Research , vol. 13, no. Feb, pp. 281–305, Feb. 2012. [Online]. Available: http://jmlr.csail.mit.edu/papers/volume13/bergstra12a/bergstra12a.pdf
2012
Cited alongside, same era.
2012
Cited alongside, same era.
2012
Cited alongside, same era.
“Imagenet large scale visual recognition challenge 2012 (ILSVRC2012),” 2012. [Online]. Available: http://www.image-net.org/challenges/LSVRC/2012/nonpub-downloads
2012
Cited alongside, same era.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 (NIPS) , F. Pereira, C. J. C. Burges et al. , Eds. Curran Associates, Inc., 2012, pp. 1097–1105. [Online]. Available: http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf
2012
Cited alongside, same era.
2012
Cited alongside, same era.
J. Stallkamp, M. Schlipsing et al. , “Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition,” Neural Networks , no. 0, pp. –, 2012. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0893608012000457
2012
Cited alongside, same era.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Machine Learning , vol. 4, no. 2, 2012. [Online]. Available: http://www.cs.toronto.edu/˜tijmen/csc321/slides/lecture_slides_lec6.pdf
2012
Cited alongside, same era.
H. Xiao, H. Xiao, and C. Eckert, “Adversarial label flips attack on support vector machines.” in ECAI , 2012, pp. 870–875. [Online]. Available: https://www.sec.in.tum.de/assets/Uploads/ecai2.pdf
2012
Cited alongside, same era.
2012
Cited alongside, same era.
I. J. Goodfellow, D. Warde-Farley et al. , “Maxout networks.” ICML , vol. 28, no. 3, pp. 1319–1327, 2013. [Online]. Available: http://www.jmlr.org/proceedings/papers/v28/goodfellow13.pdf
2013
Cited alongside, same era.
2013
Cited alongside, same era.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2016
Later among the works it cites.
M. Andrychowicz, M. Denil et al. , “Learning to learn by gradient descent by gradient descent,” in Advances in Neural Information Processing Systems 29 (NIPS) , D. D. Lee, M. Sugiyama et al. , Eds. Curran Associates, Inc., Mar. 2016, pp. 3981–3989. [Online]. Available: http://papers.nips.cc/paper/6461-learning-to-learn-by-gradient-descent-by-gradient-descent.pdf
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
A. Ng, “Nuts and bolts of building ai applications using deep learning,” NIPS Talk, Dec. 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
“MNIST for ML beginners,” Dec. 2016. [Online]. Available: https://www.tensorflow.org/tutorials/mnist/beginners/
2016
Later among the works it cites.
“tf.nn.dropout,” Dec. 2016. [Online]. Available: https://www.tensorflow.org/api_docs/python/nn/activation_functions_#dropout
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
S. Zhai, Y. Cheng et al. , “Doubly convolutional neural networks,” in Advances in Neural Information Processing Systems 29 (NIPS) , D. D. Lee, M. Sugiyama et al. , Eds. Curran Associates, Inc., Oct. 2016, pp. 1082–1090. [Online]. Available: http://papers.nips.cc/paper/6340-doubly-convolutional-neural-networks.pdf
2016
Later among the works it cites.
B. Zhou, “Places2 download,” 2016. [Online]. Available: http://places2.csail.mit.edu/download.html
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
“Kaggle cats and dogs dataset,” Oct. 2017. [Online]. Available: https://www.microsoft.com/en-us/download/details.aspx?id=54765
2017
Closest in time.
2017
Closest in time.
“Noise layers,” Jan. 2017. [Online]. Available: http://lasagne.readthedocs.io/en/latest/modules/layers/noise.html#lasagne.layers.DropoutLayer
2017
Closest in time.
2017
Closest in time.
S. Majumdar, “Densenet,” GitHub, Feb. 2017. [Online]. Available: https://github.com/titu1994/DenseNet
2017
Closest in time.
2017
Closest in time.
M. Thoma, “Master thesis (blog post),” Apr. 2017. [Online]. Available: https://martin-thoma.com/msthesis
2017
Closest in time.
2017
Closest in time.
U. Bodenhausen and S. Manke, Automatically Structured Neural Networks For Handwritten Character And Word Recognition . London: Springer London, Sep. 1993, pp. 956–961. [Online]. Available: http://dx.doi.org/10.1007/978-1-4471-2063-6_283
2063
Closest in time.