Fetching the paper…
Reading the bibliography…
The CTC-based automatic speech recognition (ASR) models without the external language model usually lack the capacity to model conditional dependencies and textual interactions.
“Connectionist temporal classification : labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Enhancing the ted-lium corpus with selected data for language modeling and more ted talks.,”
A. Rousseau, P. Deléglise, Y. Esteve, et al., · 2014
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition,”
A. Hannun, C. Case, J. Casper, B. Catanzaro, et al., · 2014
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le V, and Q. Vinyals, · 2015
Earlier work this paper cites.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, et al., · 2015
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, et al., · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J Du, X Na, B Wu, and H Zheng, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, et al., · 2017
Cited alongside, same era.
“Exploring neural transducers for end-to-end speech recognition,”
E. Battenbergand J. Chen, R. Child, et al., · 2017
Cited alongside, same era.
“Multi-accent speech recognition with hierarchical grapheme based models,”
K. Rao and H. Sak, · 2017
Cited alongside, same era.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
“Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,”
S. Kriman, S. Beliaev, B. Ginsburg, et al., · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C. Chiu, et al., · 2020
Later among the works it cites.
S. Majumdar, J. Balam, O. Hrinchuk, et al., · 2021
Later among the works it cites.
“Intermediate loss regularization for ctc-based speech recognition,”
J. Lee and S. Watanabe, · 2021
Later among the works it cites.
J. Nozaki and T. Komatsu, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Hierarchical multi task learning with ctc,”
R. Sanabria and Metze F, · 2018
Cited alongside, same era.
“Hierarchical multitask learning for ctc-based speech recognition,”
K. Krishna, S. Toshniwal, and K. Livescu, · 2018
Cited alongside, same era.
“Espnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, et al., · 2018
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, et al., · 2019
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Q. Zhang, H. Lu, H. Sak, et al., · 2020
Cited alongside, same era.
“aidatatang_200zh, a free chinese mandarin speech corpus,”
Beijing DataTang Technology Co. Lid,
Cited in the paper.
Later among the works it cites.
“Recent developments on espnet toolkit boosted by conformer,”
P. Guo, F. Boyer, X. Chang, et al., · 2021
Later among the works it cites.
“Consistent training and decoding for end-to-end speech recognition using lattice-free mmi,”
J. Tian, J. Yu, C. Weng, et al., · 2021
Later among the works it cites.
“A comparative study on non-autoregressive modelings for speech-to-text generation,”
Y. Higuchi, N. Chen, Y. Fujita, et al., · 2021
Later among the works it cites.
“Improved mask-ctc for non-autoregressive end-to-end asr,”
Y. Higuchi, H. Inaguma, S. Watanabe, et al., · 2021
Later among the works it cites.
“Hierarchical conditional end-to-end asr with ctc and multi-granular subword units,”
Y. Higuchi, K. Karube, T. Ogawa, et al., · 2022
Closest in time.
“Multi-sequence intermediate conditioning for ctc-based asr,”
Y. Fujita, T. Komatsu, and Y. Kida, · 2022
Closest in time.