Fetching the paper…
Reading the bibliography…
Commonly used automatic speech recognition (ASR) systems can be classified into frame-synchronous and label-synchronous categories, based on whether the speech is decoded on a per-frame or per-label basis.
L. Baum and J. Egon, “An in equality with applications to statistical estimation for probabilistic functions of a markov process and to a model for ecology,” Bulletin of the American Meteorological Society , vol. 73, pp. 360–363, 1967
1967
Earlier work this paper cites.
Y. Bengio and P. Frasconi, “Creadit assignment through time: Alternatives to backpropagation,” in Proc. NIPS , 1993
1993
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist speech recognition: a hybrid approach . Springer Science Business Media, 1994, vol. 247
1994
Earlier work this paper cites.
F. Jelinek, Statistical Methods for Speech Recognition . MIT Press, 1997
1997
Earlier work this paper cites.
J. G. Fiscus, “A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (ROVER),” in Proc. IEEE Workshop Autom. Speech Recognit. Understanding , 1997
1997
Earlier work this paper cites.
G. Evermann and P. C. Woodland, “Posterior probability decoding, confidence estimation and system combination,” in Speech Transcription Workshop , 2000
2000
Earlier work this paper cites.
N. Hansen and A. Ostermeier, “Completely derandomized self-adaptation in evolution strategies,” Evolutionary Computation , vol. 9, pp. 159–195, 2001
2001
Earlier work this paper cites.
J. Glass, “A probabilistic framework for segment-based speech recognition,” Computer Speech and Language , vol. 17, p. 137–152, 2003
2003
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. L. Masson, I. McCowan, W. Post, D. Reidsma, and P. Wellner, “The AMI meeting corpus: A pre-announcement,” in Proc. Int. Workshop Mach. Learn. Multimodal Interact. , 2005
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. Int. Conf. Mach. Learn. , 2006
2006
Earlier work this paper cites.
C.-H. Lee, M. Clements, S. Dusan, E. Fosler-Lussier, K. Johnson, B. Juang, and L. Rabiner, “An overview on automatic speech attribute transcription (ASAT),” in Proc. Interspeech , 2007
2007
Earlier work this paper cites.
G. Zweig and P. Nguyen, “A segmental CRF approach to large vocabulary continuous speech recognition,” in Proc. IEEE Workshop Autom. Speech Recognit. Understanding , 2009
2009
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlícek, Y. Qian, P. Schwarz, J. Silovský, G. Stemmer, and K. Veselý, “The Kaldi speech recognition toolkit,” in Proc. IEEE Workshop Autom. Speech Recognit. Understanding , 2011
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in ICML Workshop on Representation Learn. , 2012
2012
Earlier work this paper cites.
H. Su, G. Li, D. Yu, and F. Seide, “Error back propagation for sequence training of context-dependent deep networks for conversational speech transcription,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2013
2013
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proc. Int. Conf. Learn. Representations , 2015
2015
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Adv. Neural Inf. Proces. Syst. , 2015
2015
Earlier work this paper cites.
L. Lu, X. Zhang, K. Cho, and S. Renals, “A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition,” in Proc. Interspeech , 2015
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2016
2016
Cited alongside, same era.
G. Pundak and T. N. Sainath, “Lower frame rate neural network acoustic models,” in Proc. Interspeech , 2016
2016
Cited alongside, same era.
X. Liu, X. Chen, Y. Wang, M. J. F. Gales, and P. C. Woodland, “Two efficient lattice rescoring methods using recurrent neural network language models,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 24, pp. 1438–1449, 2016
2016
Cited alongside, same era.
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. Interspeech , 2019
2019
Later among the works it cites.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Training language models for long-span cross-sentence evaluation,” in Proc. IEEE Workshop Autom. Speech Recognit. Understanding , 2019
2019
Later among the works it cites.
M. Kitza, P. Golik, R. Schlüter, and H. Ney, “Cumulative adaptation for BLSTM acoustic models,” in Proc. Interspeech , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen, X. Liu, Y. Wang, M. Gales, and P. Woodland, “Efficient training and evaluation of recurrent neural network language models for automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, pp. 2146–2157, 2016
2016
Cited alongside, same era.
R. Prabhavalkar, K. Rao, T. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in Proc. Interspeech , 2017
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE J. Sel. Topics Signal Process. , vol. 11, pp. 1240–1253, 2017
2017
Cited alongside, same era.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inf. Proces. Syst. , 2017
2017
Cited alongside, same era.
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “End-to-end speech recognition using lattice-free MMI,” in Proc. Interspeech , 2018
2018
Cited alongside, same era.
D. Povey, G. Cheng, Y. Wang, K. Li, H. Xu, M. Yarmohammadi, and S. Khudanpur, “Semi-orthogonal low-rank matrix factorization for deep neural networks,” in Proc. Interspeech , 2018
2018
Cited alongside, same era.
C. Chiu, T. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, “State-of-the-art speech recognition with sequence-to-sequence models,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2018
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech , 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2020
2020
Later among the works it cites.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for speech recognition,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
H. Miao, G. Cheng, P. Zhang, and Y. Yan, “Online hybrid CTC/attention end-to-end automatic speech recognition architecture,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 28, pp. 1452–1465, 2020
2020
Later among the works it cites.
N. Moritz, T. Hori, and J. L. Roux, “Streaming automatic speech recognition with the Transformer model,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2020
2020
Later among the works it cites.
W. Wang, Y. Zhou, C. Xiong, and R. Socher, “An investigation of phone-based subword units for end-to-end speech recognition,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard-300,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
D. Jiang, C. Zhang, and P. Woodland, “Variable frame rate acoustic models usingminimum error reinforcement learning,” in Proc. Interspeech , 2021
2021
Closest in time.
R. Prabhavalkar, Y. He, D. Rybach, S. Campbell, A. Narayanan, T. Strohman, and T. Sainath, “Less is more: Improved RNN-T decoding using limited label context and path merging,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2021
2021
Closest in time.
G. Sun, C. Zhang, and P. C. Woodland, “Transformer language models with LSTM-based cross-utterance information representation,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. , 2021
2021
Closest in time.
2021
Closest in time.