Fetching the paper…
Reading the bibliography…
An embedding-based speaker adaptive training (SAT) approach is proposed and investigated in this paper for deep neural network acoustic modeling.
J. L. Gauvain and C. H. Lee, “Maximum a posteriori estimation for multivariate gaussian mixture observations of markov chains,”
1994
Earlier work this paper cites.
C. J. Leggetter and P. C. Woodland, “Maximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models,”
1995
Earlier work this paper cites.
T. Anastasakos, J. McDonough, R. Schwartz, and J. Makhoul, “A compact model for speaker-adaptive training,” in
1996
Earlier work this paper cites.
M. J. F. Gales, “Maximum likelihood linear transformations for HMM-based speech recognition,”
1998
Earlier work this paper cites.
H. Shimodaira, “Improving predictive inference under covariate shift by weighting the log-likelihood function,”
2000
Earlier work this paper cites.
M. J. F. Gales, “Cluster adaptive training of hidden Markov models,”
2000
Earlier work this paper cites.
B. Li and K. C. Sim, “Comparison of discriminative input and output transformations for speaker adaptation in the hybrid NN/HMM systems,” in
2010
Cited alongside, same era.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Cited alongside, same era.
K. Yao, D. Yu, F. Seide, H. Su, L. Deng, and Y. Gong, “Adaptation of context-dependent deep neural networks for automatic speech recognition,” in
2012
Cited alongside, same era.
H. Liao, “Speaker adaptation of context dependent deep neural networks,” in
2013
Cited alongside, same era.
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker adaptation of neural network acoustic models using I-vectors,” in
2013
Cited alongside, same era.
P. Swietojanski and S. Renals, “Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,” in
2014
Later among the works it cites.
Y. Miao, H. Zhang, and F. Metze, “Speaker adaptive training of deep neural network acoustic models using I-vectors,”
2015
Later among the works it cites.
G. Saon, T. Sercu, S. Rennie, and H.-K. Kuo, “The IBM 2016 English conversational telephone speech recognition system,” in
2016
Later among the works it cites.
P. Swietojanski and S. Renals, “SAT-LHUC: speaker adaptive training for learning hidden unit contributions,” in
2016
Later among the works it cites.
L. Samarakoon and K. C. Sim, “Factorized hidden layer adaptation for deep neural network based acoustic modeling,”
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Xue, J. Li, D. Yu, M. Seltzer, and Y. Gong, “Singular value decomposition based low-footprint speaker adaptation and personalization for deep neural network,” in
2014
Cited alongside, same era.
Https://www.iarpa.gov/index.php/research-programs/babel
Cited in the paper.
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim, B. Roomi, and P. Hall, “English conversational telephone speech recognition by humans and machines,” in
2017
Closest in time.