T. Ash, “Dynamic node creation in backpropagation networks,” Connection Science , vol. 1, no. 4, pp. 365–375, 1989
1989
Earlier work this paper cites.
M. Mezard and J.-P. Nadal, “Learning in feedforward layered networks: The tiling algorithm,” J. of Physics A: Mathematical and General , vol. 22, no. 12, p. 2191, 1989
1989
Earlier work this paper cites.
S. Lowel and W. Singer, “Selection of intrinsic horizontal connections in the visual cortex by correlated neuronal activity,” Science , vol. 255, no. 5041, p. 209, 1992
1992
Earlier work this paper cites.
C. J. Burges and B. Schölkopf, “Improving the accuracy and speed of support vector machines,” in Proc. Advances in Neural Information Processing Systems , 1997, pp. 375–381
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
K. O. Stanley and R. Miikkulainen, “Evolving neural networks through augmenting topologies,” Evolutionary Computation , vol. 10, no. 2, pp. 99–127, 2002
2002
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li, “ImageNet: A large-scale hierarchical image database,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in Proc. Advances in Neural Information Processing Systems , vol. 2011, no. 2, 2011, p. 5
2011
Earlier work this paper cites.
D. Yu, F. Seide, G. Li, and L. Deng, “Exploiting sparseness in deep neural networks for large vocabulary speech recognition,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing , 2012, pp. 4409–4412
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Proc. Advances in Neural Information Processing Systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
A. Graves, A.-R. Mohamed, and G. E. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing , 2013, pp. 6645–6649
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Proc. Advances in Neural Information Processing Systems , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
Y. Gong, L. Liu, M. Yang, and L. Bourdev, “Compressing deep convolutional networks using vector quantization,” arXiv preprint arXiv:1412.6115 , 2014
Original
2014
Earlier work this paper cites.
A. Krizhevsky, “One weird trick for parallelizing convolutional neural networks,” arXiv preprint arXiv:1404.5997 , 2014
Original
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proc. ACM Int. Conf. Multimedia , 2014, pp. 675–678
2014
Earlier work this paper cites.
M. D. Collins and P. Kohli, “Memory bounded deep convolutional networks,” arXiv preprint arXiv:1412.1442 , 2014
Original
2014
Earlier work this paper cites.