Fetching the paper…
Reading the bibliography…
Sequence discriminative training criteria have long been a standard tool in automatic speech recognition for improving the performance of acoustic models over their maximum likelihood / cross entropy trained counterparts.
H. A. Bourlard and N. Morgan, “Connectionist speech recognition: a hybrid approach,” vol. 247, 1994
1994
Earlier work this paper cites.
Y. Normandin, “Maximum mutual information estimation of hidden markov models,”
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
J. Binder, K. P. Murphy, and S. J. Russell, “Space-efficient inference in dynamic probabilistic networks,” in
1997
Earlier work this paper cites.
R. Schlüter, B. Müller, F. Wessel, and H. Ney, “Interdependence of language models and discriminative training,” in
1999
Earlier work this paper cites.
G. Zweig and M. Padmanabhan, “Exact alpha-beta computation in logarithmic space with application to MAP word graph construction,” in
2000
Earlier work this paper cites.
D. Povey, “Discriminative training for large vocabulary speech recognition,” Ph.D. dissertation, Cambridge University, 01 2003
2003
Earlier work this paper cites.
M. Gibson and T. Hain, “Hypothesis spaces for minimum bayes risk training in large vocabulary speech recognition,” in
2006
Earlier work this paper cites.
S. Chen, B. Kingsbury, L. Mangu, D. Povey, G. Saon, H. Soltau, and G. Zweig, “Advances in speech transcription at IBM under the DARPA EARS program,”
2006
Cited alongside, same era.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks,” in
2011
Cited alongside, same era.
B. Kingsbury, T. N. Sainath, and H. Soltau, “Scalable minimum bayes risk training of deep neural network acoustic models using distributed hessian-free optimization,” pp. 10–13, 2012
2012
Cited alongside, same era.
K. Veselý, A. Ghoshal, L. Burget, and D. Povey, “Sequence-discriminative training of deep neural networks,” in
2013
Cited alongside, same era.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,”
2014
N. Kanda, Y. Fujita, and K. Nagamatsu, “Investigation of lattice-free maximum mutual information-based acoustic models with sequence-level kullback-leibler divergence,” in
2017
Later among the works it cites.
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” pp. 2999–3007, 2017
2017
Later among the works it cites.
A. Zeyer, E. Beck, R. Schlüter, and H. Ney, “CTC in the context of generalized full-sum HMM training,” in
2017
Later among the works it cites.
N. Kanda, Y. Fujita, and K. Nagamatsu, “Lattice-free state-level minimum bayes risk training of acoustic models,” in
2018
Later among the works it cites.
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, “The microsoft 2017 conversational speech recognition system,” Tech. Rep., 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
P. Voigtlaender, P. Doetsch, S. Wiesler, R. Schlüter, and H. Ney, “Sequence-discriminative training of recurrent neural networks,” in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in
2016
Cited alongside, same era.
“Quaero.” [Online]. Available:
Cited in the paper.
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “Flat-start single-stage discriminatively trained hmm-based models for ASR,”
2018
Later among the works it cites.
H. Hadian, D. Povey, H. Sameti, J. Trmal, and S. Khudanpur, “Improving LF-MMI using unconstrained supervisions for ASR,” in
2018
Later among the works it cites.