Fetching the paper…
Reading the bibliography…
We describe the neural-network training framework used in the Kaldi speech recognition toolkit, which is geared towards training DNNs with large amounts of training data using multiple GPU-equipped or multi-core machines.
Some methods of speeding up the convergence of iteration methods
Polyak, B.T · 1964
Earlier work this paper cites.
Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
Davis, Steven and Mermelstein, Paul · 1980
Earlier work this paper cites.
Maximum mutual information estimation of hidden markov model parameters for speech recognition
Bahl, L Brown, de Souza, P, and P Mercer, R · 1986
Earlier work this paper cites.
Mean and Variance Adaptation Within the MLLR Framework
Gales, M. J. F. and Woodland, P. C · 1996
Earlier work this paper cites.
Estimating an eigenvector by the power method with a random start
Del Corso, Gianna M · 1997
Earlier work this paper cites.
Natural gradient descent for training multi-layer perceptrons
Yang, Howard Hua and Amari, Shun-ichi · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, Shun-Ichi · 1998
Earlier work this paper cites.
Online learning and stochastic approximations
Bottou, Léon · 1998
Earlier work this paper cites.
Complexity issues in natural gradient descent method for training multilayer perceptrons
Yang, Howard Hua and Amari, Shun-ichi · 1998
Earlier work this paper cites.
Statistical analysis of learning dynamics
Murata, Noboru and Amari, Shun-ichi · 1999
Earlier work this paper cites.
Minimum Phone Error and I-smoothing for Improved Discriminative Training
Povey., D. and Woodland, P. C · 2002
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Kushner, Harold J and Yin, George · 2003
Earlier work this paper cites.
Discriminative Training for Large Voculabulary Speech Recognition
Povey, D · 2004
Cited alongside, same era.
A fast learning algorithm for deep belief nets
Hinton, Geoffrey, Osindero, Simon, and Teh, Yee-Whye · 2006
Cited alongside, same era.
Hypothesis Spaces For Minimum Bayes Risk Training In Large Vocabulary Speech Recognition
M., Gibson and T., Hain · 2006
Cited alongside, same era.
Adaptive online gradient descent
Hazan, Elad, Rakhlin, Alexander, and Bartlett, Peter L · 2007
Cited alongside, same era.
Evaluation of proposed modifications to MPE for large scale discriminative training
Povey, Daniel and Kingsbury, Brian · 2007
Cited alongside, same era.
Topmoumoute online natural gradient algorithm
Roux, Nicolas Le, Bengio, Yoshua, and antoine Manzagol, Pierre · 2007
Cited alongside, same era.
Feature engineering in context-dependent deep neural networks for conversational speech transcription
Seide, Frank, Li, Gang, Chen, Xie, and Yu, Dong · 2011
Later among the works it cites.
Large Scale Distributed Deep Networks
Dean, Jeffrey, Corrado, Greg S., Monga, Rajat, Chen, Kai, Devin, Matthieu, Le, Quoc V., Mao, Mark Z., Ranzato, Marc’Aurelio, Senior, Andrew, Tucker, Paul, Yang, Ke, and Ng, Andrew Y · 2012
Later among the works it cites.
Efficient backprop
LeCun, Yann A, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 2012
Later among the works it cites.
Rapid adaptation for mobile speech applications
Bacchiani, Michiel · 2013
Later among the works it cites.
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Later among the works it cites.
Rectifier nonlinearities improve neural network acoustic models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boosted MMI for Feature and Model Space Discriminative Training
Povey, D., Kanevsky, D., Kingsbury, B., Ramabhadran, B., Saon, G., and Visweswariah, K · 2008
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Cited alongside, same era.
Investigation of full-sequence training of deep belief networks for speech recognition
Mohamed, Abdel-rahman, Yu, Dong, and Deng, Li · 2010
Cited alongside, same era.
A simplified natural gradient learning algorithm
Bastian, Michael R, Gunther, Jacob H, and Moon, Todd K · 2011
Cited alongside, same era.
Front-end factor analysis for speaker verification
Dehak, Najim, Kenny, Patrick, Dehak, Réda, Dumouchel, Pierre, and Ouellet, Pierre · 2011
Cited alongside, same era.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Niu, Feng, Recht, Benjamin, Ré, Christopher, and Wright, Stephen J · 2011
Cited alongside, same era.
Maas, Andrew L, Hannun, Awni Y, and Ng, Andrew Y · 2013
Later among the works it cites.
Pascanu, Razvan and Bengio, Yoshua · 2013
Later among the works it cites.
Speaker adaptation of neural network acoustic models using i-vectors
Saon, George, Soltau, Hagen, Nahamoo, David, and Picheny, Michael · 2013
Later among the works it cites.
An empirical study of learning rates in deep neural networks for speech recognition
Senior, Andrew, Heigold, Georg, Ranzato, Marc’Aurelio, and Yang, Ke · 2013
Later among the works it cites.
Sequence-discriminative training of deep neural networks
Veselỳ, Karel, Ghoshal, Arnab, Burget, Lukáš, and Povey, Daniel · 2013
Later among the works it cites.
Improving dnn speaker independence with i-vector inputs
Senior, Andrew and Lopez-Moreno, Ignacio · 2014
Closest in time.
Improving deep neural network acoustic models using generalized maxout networks
Zhang, Xiaohui, Trmal, Jan, Povey, Daniel, and Khudanpur, Sanjeev · 2014
Closest in time.