Low-rank Matrix Factorization for Deep Neural Network Training with High-dimensional Output Targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arısoy, and Bhuvana Ramabhadran · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton · 2013
Later among the works it cites.
Sequence-discriminative training of deep neural networks
Karel Veselý, Arnab Ghoshal, Lukás Burget, and Daniel Povey · 2013
Later among the works it cites.
Convolutional neural networks for speech recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu · 2014
Later among the works it cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Later among the works it cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Monica Bianchini and Franco Scarselli · 2014
Later among the works it cites.
Scalable kernel methods via doubly stochastic gradients
Bo Dai, Bo Xie, Niao He, Yingyu Liang, Anant Raj, Maria-Florina Balcan, and Le Song · 2014
Later among the works it cites.
Compact random feature maps
Raffay Hamid, Ying Xiao, Alex Gittens, and Dennis DeCoste · 2014
Later among the works it cites.
Kernel methods match deep neural networks on TIMIT
Po-Sen Huang, Haim Avron, Tara N. Sainath, Vikas Sindhwani, and Bhuvana Ramabhadran · 2014
Later among the works it cites.
A simple proof of Stirling’s formula for the gamma function
G. J. O. Jameson · 2014
Later among the works it cites.
On the number of linear regions of deep neural networks
Guido F. Montúfar, Razvan Pascanu, KyungHyun Cho, and Yoshua Bengio · 2014
Later among the works it cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Hasim Sak, Andrew W. Senior, and Françoise Beaufays · 2014
Later among the works it cites.
Personal communication, 2014
Alex Smola · 2014
Later among the works it cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Sparse random feature algorithm as coordinate descent in Hilbert space
E.-H. Yen, T.-W. Lin, S.-D. Lin, P.K. Ravikumar, and I.S. Dhillon · 2014
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Spherical random features for polynomial kernels
Jeffrey Pennington, Felix X. Yu, and Sanjiv Kumar · 2015
Later among the works it cites.
Compact nonlinear maps and circulant extensions
Original
Felix X. Yu, Sanjiv Kumar, Henry A. Rowley, and Shih-Fu Chang · 2015
Later among the works it cites.
A comparison between deep neural nets and kernel acoustic models for speech recognition
Zhiyun Lu, Dong Quo, Alireza Bagheri Garakani, Kuan Liu, Avner May, Aurélien Bellet, Linxi Fan, Michael Collins, Brian Kingsbury, Michael Picheny, and Fei Sha · 2016
Later among the works it cites.
Compact kernel models for acoustic modeling via random feature selection
Avner May, Michael Collins, Daniel J. Hsu, and Brian Kingsbury · 2016
Later among the works it cites.
Purely sequence-trained neural networks for asr based on lattice-free mmi
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahrmani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur · 2016
Later among the works it cites.
Achieving human parity in conversational speech recognition
Original
W. Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig · 2016
Later among the works it cites.