Fetching the paper…
Reading the bibliography…
Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
“AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Joint CTC/attention decoding for end-to-end speech recognition,”
T. Hori, S. Watanabe, and J. Hershey, · 2017
Earlier work this paper cites.
“Acoustic-to-word attention-based model complemented with character-level ctc-based model,”
S. Ueno, H. Inaguma, M. Mimura, and T. Kawahara, · 2018
Earlier work this paper cites.
“AISHELL-2: transforming mandarin ASR research into industrial scale,”
J. Du, X. Na, X. Liu, and H. Bu, · 2018
Earlier work this paper cites.
“Espnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, and N. Chen, · 2018
Earlier work this paper cites.
“The speechtransformer for large-scale mandarin chinese speech recognition,”
Y. zhao, J. Li, X. Wang, and Y. Li, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. Chang, K. Lee, and K. Toutanova, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, · 2019
Earlier work this paper cites.
“fairseq: A fast, extensible toolkit for sequence modeling,”
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, · 2019
Cited alongside, same era.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,”
D. S. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
A. Baevski, S. Schneider, and M. Auli, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict,”
Y. Higuchi, S. Watanabe, N. Chen, T. Ogawa, and T. Kobayashi, · 2020
Cited alongside, same era.
“Imputer: Sequence modelling via imputation and dynamic programming,”
“Transformers: State-of-the-art natural language processing,”
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. v. Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, · 2020
Later among the works it cites.
“Intermediate loss regularization for ctc-based speech recognition,”
J. Lee and S. Watanabe, · 2021
Later among the works it cites.
“Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition,”
C. Yi, S. Zhou, and B. Xu, · 2021
Later among the works it cites.
“Non-autoregressive transformer-based end-to-end ASR using BERT,”
F. Yu and K. Chen, · 2021
Later among the works it cites.
“Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-Supervised Learning,”
K. Deng, S. Cao, and L. Ma, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Chan, C. Saharia, G. Hinton, M. Norouzi, and N. Jaitly, · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, · 2020
Cited alongside, same era.
“Distilling the knowledge of bert for sequence-to-sequence asr,”
H. Futami, H. Inaguma, S. Ueno, M. Mimura, S. Sakai, and T. Kawahara, · 2020
Cited alongside, same era.
“Sequence-to-sequence automatic speech recognition with word embedding regularization and fused decoding,”
A. H. Liu, T. Sung, S. Chuang, H. Lee, and L. Lee, · 2020
Cited alongside, same era.
“Adapt-and-adjust: Overcoming the long-tail problem of multilingual speech recognition,”
G. I. Winata, G. Wang, C. Xiong, and S. C. H. Hoi, · 2020
Cited alongside, same era.
“Cif: Continuous integrate-and-fire for end-to-end speech recognition,”
L. Dong and B. Xu, · 2020
Cited alongside, same era.
“Semantics of the unwritten: The effect of end of paragraph and sequence tokens on text generation with GPT2,”
H. Bai, P. Shi, J. Lin, L. Tan, K. Xiong, W. Gao, J. Liu, and M. Li, · 2021
Later among the works it cites.
“Fast end-to-end speech recognition via non-autoregressive models and cross-modal knowledge transferring from bert,”
Y. Bai, J. Yi, J. Tao, Z. Tian, Z. Wen, and S. Zhang, · 2021
Later among the works it cites.
“Improving hybrid ctc/attention end-to-end speech recognition with pretrained acoustic and language models,”
K. Deng, S. Cao, Y. Zhang, and L. Ma, · 2021
Later among the works it cites.
“An Improved Single Step Non-Autoregressive Transformer for Automatic Speech Recognition,”
R. Fan, W. Chu, P. Chang, J. Xiao, and A. Alwan, · 2021
Later among the works it cites.
“Alleviating asr long-tailed problem by decoupling the learning of representation and classification,”
K. Deng, G. Cheng, R. Yang, and Y. Yan, · 2022
Closest in time.