Fetching the paper…
Reading the bibliography…
While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process.
“SWITCHBOARD: Telephone speech corpus for research and development,”
J. J. Godfrey, E. C. Holliman, and J. McDaniel, · 1992
Earlier work this paper cites.
“A new algorithm for data compression,”
P. Gage, · 1994
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Earlier work this paper cites.
“Joint CTC/attention decoding for end-to-end speech recognition,”
T. Hori, S. Watanabe, and J. Hershey, · 2017
Earlier work this paper cites.
“AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, · 2017
Earlier work this paper cites.
“AISHELL-2: Transforming mandarin ASR research into industrial scale,”
J. Du, X. Na, X. Liu, and H. Bu, · 2018
Earlier work this paper cites.
“ESPnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, and N. Chen, · 2018
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. Chang, K. Lee, and K. Toutanova, · 2019
Earlier work this paper cites.
“RoBERTa: A robustly optimized BERT pretraining approach,”
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, · 2019
Cited alongside, same era.
“Spike-triggered non-autoregressive Transformer for end-to-end speech recognition,”
Z. Tian, J. Yi, J. Tao, Y. Bai, S. Zhang, and Z. Wen, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“CIF: Continuous integrate-and-fire for end-to-end speech recognition,”
L. Dong and B. Xu, · 2020
Cited alongside, same era.
“TSNAT: Two-step non-autoregressvie transformer models for speech recognition,”
Z. Tian, J. Yi, J. Tao, Y. Bai, S. Zhang, Z. Wen, and X. Liu, · 2021
Later among the works it cites.
“Improving accent identification and accented speech recognition under a framework of self-supervised learning,”
K. Deng, S. Cao, and L. Ma, · 2021
Later among the works it cites.
“Non-autoregressive transformer-based end-to-end ASR using BERT,”
F. Yu and K. Chen, · 2021
Later among the works it cites.
“Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition,”
C. Yi, S. Zhou, and B. Xu, · 2021
Later among the works it cites.
“Speech recognition by simply fine-tuning BERT,”
W. Huang, C. Wu, S. Luo, K. Chen, H. Wang, and T. Toda, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Dong, C. Yi, J. Wang, S. Zhou, S. Xu, X. Jia, and B. Xu, · 2020
Cited alongside, same era.
“Transformers: State-of-the-art natural language processing,”
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. v. Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, · 2020
Cited alongside, same era.
“Applying wav2vec2.0 to speech recognition in various low-resource languages,”
C. Yi, J. Wang, N. Cheng, S. Zhou, and B. Xu, · 2020
Cited alongside, same era.
“Improved mask-CTC for non-autoregressive end-to-end ASR,”
Y. Higuchi, H. Inaguma, S. Watanabe, T. Ogawa, and T. Kobayashi, · 2021
Cited alongside, same era.
“Non-autoregressive Transformer ASR with CTC-enhanced decoder input,”
X. Song, Z. Wu, Y. Huang, C. Weng, D. Su, and H. Meng, · 2021
Cited alongside, same era.
“Non-autoregressive Transformer for speech recognition,”
N. Chen, S. Watanabe, J. Villalba, P. Zelasko, and N. Dehak, · 2021
Cited alongside, same era.
K. Deng, S. Cao, Y. Zhang, and L. Ma, · 2021
Later among the works it cites.
“Non-autoregressive deliberation-attention based end-to-end ASR,”
C. Gao, G. Cheng, J. Zhou, P. Zhang, and Y. Yan, · 2021
Later among the works it cites.
“CASS-NAT: CTC alignment-based single step non-autoregressive transformer for speech recognition,”
R. Fan, W. Chu, P. Chang, and J. Xiao, · 2021
Later among the works it cites.
“Fast end-to-end speech recognition via non-autoregressive models and cross-modal knowledge transferring from BERT,”
Y. Bai, J. Yi, J. Tao, Z. Tian, Z. Wen, and S. Zhang, · 2021
Later among the works it cites.
“Alleviating asr long-tailed problem by decoupling the learning of representation and classification,”
K. Deng, G. Cheng, R. Yang, and Y. Yan, · 2022
Closest in time.
“Transformer-based end-to-end speech recognition with residual gaussian-based self-attention,”
C. Liang, M. Xu, and X.-L. Zhang, · 2076
Closest in time.