Fetching the paper…
Reading the bibliography…
This paper presents novel Weighted Finite-State Transducer (WFST) topologies to implement Connectionist Temporal Classification (CTC)-like algorithms for automatic speech recognition.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006
2006
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep rnn models and wfst-based decoding,” in ASRU , 2015
2015
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, O. İrsoy, A. Graves, F. Beaufays, and J. Schalkwyk, “Learning acoustic frame labeling for speech recognition with recurrent neural networks,” in ICASSP , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI,” in Interspeech , 2016
2016
Earlier work this paper cites.
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “End-to-end Speech Recognition Using Lattice-free MMI,” in Interspeech , 2018
2018
Earlier work this paper cites.
H. Liu, S. Jin, and C. Zhang, “Connectionist temporal classification with maximum entropy regularization,” in NeuRIPS , 2018
2018
Earlier work this paper cites.
2019
Cited alongside, same era.
H. Xiang and Z. Ou, “CRF-based single-stage acoustic modeling with CTC topology,” in ICASSP , 2019
2019
Cited alongside, same era.
D. Le, X. Zhang, W. Zheng, C. Fügen, G. Zweig, and M. L. Seltzer, “From senones to chenones: Tied context-dependent graphemes for hybrid speech recognition,” in ASRU , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
F. Zhang, Y. Wang, X. Zhang, C. Liu, Y.Saraf, and G. Zweig, “Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces,” in Interspeech , 2020
Y. Shao, Y. Wang, D. Povey, and S. Khudanpur, “PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR,” in Interspeech , 2020
2020
Later among the works it cites.
X. Zhang, F. Zhang, C. Liu, K. Schubert, J. Chan, P. Prakash, J. Liu, C. Yeh, F. Peng, Y. Saraf, and G. Zweig, “Benchmarking LF-MMI, CTC and RNN-T criteria for streaming ASR,” in SLT , 2021
2021
Closest in time.
2021
Closest in time.
D. Povey, P. Żelasko, and S. Khudanpur, “Speech recognition with next-generation kaldi (k2, lhotse, icefall),” Interspeech: tutorials , 2021
2021
Closest in time.
A. Zeyer, R. Schlüter, and H. Ney, “Why does CTC result in peaky behavior?” arXiv:2105.14849 , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
L. Hongzhu and W. Weiqiang, “Reinterpreting CTC training as iterative fitting,” Pattern Recognition. , 2020
2020
Cited alongside, same era.
H. Braun, J. Luitjens, R. Leary, T. Kaldewey, and D. Povey, “GPU-accelerated Viterbi exact lattice decoder for batched online and offline speech recognition,” in ICASSP , 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N.Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Interspeech , 2020
2020
Cited alongside, same era.
2021
Closest in time.
N. Moritz, T. Hori, and J. L. Roux, “Semi-supervised speech recognition via graph-based temporal classification,” in ICASSP , 2021
2021
Closest in time.