Fetching the paper…
Reading the bibliography…
State-level minimum Bayes risk (sMBR) training has become the de facto standard for sequence-level training of speech recognition acoustic models.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
1992
Earlier work this paper cites.
V. Valtchev, J. Odell, P. C. Woodland, and S. J. Young, “Lattice-based discriminative training for large vocabulary speech recognition,” in
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour
1999
Earlier work this paper cites.
J. Kaiser, B. Horvat, and Z. Kacic, “A novel loss function for the overall risk criterion based discriminative training of HMM models,” in
2000
Earlier work this paper cites.
V. Goel and W. J. Byrne, “Minimum Bayes-risk automatic speech recognition,”
2000
Earlier work this paper cites.
J. Eisner, “Expectation semirings: Flexible EM for learning finite-state transducers,” in
2001
Earlier work this paper cites.
M. Mohri and M. Riley, “A weight pushing algorithm for large vocabulary speech recognition,” in
2001
Earlier work this paper cites.
D. Povey and P. C. Woodland, “Minimum phone error and I-smoothing for improved discriminative training,” in
2002
Earlier work this paper cites.
M. Mohri, F. Pereira, and M. Riley, “Weighted finite-state transducers in speech recognition,”
2002
Earlier work this paper cites.
J. Eisner, “Parameter estimation for probabilistic finite-state transducers,” in
2002
Earlier work this paper cites.
V. Doumpiotis and W. Byrne, “Pinched lattice minimum Bayes risk discriminative training for large vocabulary continuous speech recognition,” in
2004
Earlier work this paper cites.
J. Zheng and A. Stolcke, “Improved discriminative training using phone lattices,” in
2005
Cited alongside, same era.
W. Macherey, L. Haferkamp, R. Schlüter, and H. Ney, “Investigations on error minimizing training criteria for discriminative training in automatic speech recognition,” in
2005
Cited alongside, same era.
G. Heigold, W. Macherey, R. Schluter, and H. Ney, “Minimum exact word error training,” in
2005
Cited alongside, same era.
M. Gibson and T. Hain, “Hypothesis spaces for minimum Bayes risk training in large vocabulary speech recognition,” in
2006
Cited alongside, same era.
D. Povey and B. Kingsbury, “Evaluation of proposed modifications to MPE for large scale discriminative training,” in
2007
Cited alongside, same era.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2012
Later among the works it cites.
K. Veselỳ, A. Ghoshal, L. Burget, and D. Povey, “Sequence-discriminative training of deep neural networks,” in
2013
Later among the works it cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Later among the works it cites.
V. Mnih, N. Heess, A. Graves
2014
Later among the works it cites.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and accurate recurrent neural network acoustic models for speech recognition,” in
2015
Later among the works it cites.
R. C. Van Dalen and M. J. Gales, “Annotating large lattices with the exact word error,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Heigold, T. Deselaers, R. Schlüter, and H. Ney, “Modified MMI/MPE: A direct evaluation of the margin in speech recognition,” in
2008
Cited alongside, same era.
Z.-J. Yan, B. Zhu, Y. Hu, and R.-H. Wang, “Minimum word classification error training of HMMs for automatic speech recognition,” in
2008
Cited alongside, same era.
M. Gibson, “Minimum Bayes risk acoustic model estimation and adaptation,” Ph.D. dissertation, University of Sheffield, UK, 2008
2008
Cited alongside, same era.
M. Mohri, F. Pereira, and M. Riley, “Speech recognition with weighted finite-state transducers,” in
2008
Cited alongside, same era.
B. Kingsbury, “Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,” in
2009
Cited alongside, same era.
H.-A. Loeliger and M. Molkaraie, “Estimating the partition function of 2-D fields and the capacity of constrained noiseless 2-D channels using tree-based Gibbs sampling,” in
2009
Cited alongside, same era.
F. Sehnke, C. Osendorfer, T. Rückstieß, A. Graves, J. Peters, and J. Schmidhuber, “Parameter-exploring policy gradients,”
2010
Cited alongside, same era.
2015
Later among the works it cites.
2015
Later among the works it cites.
J. R. Bellegarda and C. Monz, “State of the art in statistical methods for language and speech processing,”
2016
Later among the works it cites.
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, “Sequence level training with recurrent neural networks,” in
2016
Later among the works it cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahrmani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” 2016
2016
Later among the works it cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard
2016
Later among the works it cites.