Fetching the paper…
Reading the bibliography…
Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent structures on a number of ASR tasks.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,” arXiv:1412.6980, 2014
D. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Effective approaches to attention-based neural machine translation,”
T. Luong, H. Pham, and C. D. Manning, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Earlier work this paper cites.
“Layer normalization,” arXiv:1607.06450, 2016
J. Ba, J. Kiros, and G. Hinton, · 2016
Earlier work this paper cites.
“Adaptive computation time for recurrent neural networks,” arXiv:1603.08983, 2016
A. Graves, · 2016
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Earlier work this paper cites.
“Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition,”
T. Hori, S. Watanabe, and J. R. Hershey, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Local monotonic attention mechanism for end-to-end speech and language processing,”
A. Tjandra, S. Sakti, and S. Nakamura, · 2017
Cited alongside, same era.
“Gaussian prediction based attention for online end-to-end speech recognition,”
J. Hou, S. Zhang, and L. Dai, · 2017
Cited alongside, same era.
“Online and linear-time attention by enforcing monotonic alignments,”
T. Luong, H. Pham, and C. D. Manning, · 2017
Cited alongside, same era.
“AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, · 2017
Cited alongside, same era.
“Cold fusion: Training seq2seq models together with language models,”
A. Sriram, H. Jun, S. Satheesh, and A. Coates, · 2018
“Towards online end-to-end transformer automatic speech recognition,” arXiv:1910.11871, 2019
E. Tsunoo, Y. Kashiwagi, T. Kumakura, and S. Watanabe, · 2019
Later among the works it cites.
“End-to-end speech recognition with adaptive computation steps,”
M. Li, M. Liu, and H. Masanori, · 2019
Later among the works it cites.
“ESPnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, and et al., · 2019
Later among the works it cites.
“Self-attention transducers for end-to-end speech recognition,”
Z. Tian, J. Yi, J. Tao, Y. Bai, and Z. Wen, · 2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
L. Dong, S. Xu, and B. Xu, · 2018
Cited alongside, same era.
“Monotonic chunkwise attention,”
C.-C. Chiu and C. Raffel, · 2018
Cited alongside, same era.
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
C. Shan, W. Chao, W. Guangsen, S. Dan, L. Min, Y. Dong, and L. Xie, · 2019
Cited alongside, same era.
“Self-attention aligner: A latency-control end-to-end model for ASR using self-attention network and chunk-hopping,”
L. Dong, F. Wang, and B. Xu, · 2019
Cited alongside, same era.
“The speechtransformer for large-scale mandarin chinese speech recognition,”
Y. Zhao, J. Li, X. Wang, and Y. Li, · 2019
Cited alongside, same era.
“A comparative study on Transformer vs RNN in speech applications,”
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, and W. Zhang, · 2019
Cited alongside, same era.
“Streaming attention,” https://github.com/HaoranMiao/streaming-attention, 2020,
G. Cheng H. Miao, · 2020
Closest in time.
“Transformer-based online CTC/attention end-to-end speech recognition architecture,”
H. Miao, G. Cheng, C. Gao, P. Zhang, and Y. Yan, · 2020
Closest in time.
“Streaming automatic speech recognition with the Transformer model,”
N. Moritz, T. Hori, and J. L. Roux, · 2020
Closest in time.
“Online hybrid ctc/attention end-to-end automatic speech recognition architecture,”
H. Miao, G. Cheng, P. Zhang, and Y. Yan, · 2020
Closest in time.
“CIF: Continuous integrate-and-fire for end-to-end speech recognition,”
L. Dong and B. Xu, · 2020
Closest in time.
“Synchronous transformers for end-to-end speech recognition,”
Z. Tian, J. Yi, Y. Bai, J. Tao, S. Zhang, and Z. Wen, · 2020
Closest in time.
“Multi-head monotonic chunkwise attention for online speech recognition,” arXiv:2005.00205, 2020
B. Liu, S. Cao, S. Sun, W. Zhang, and L. Ma, · 2020
Closest in time.
“Enhancing monotonic multihead attention for streaming asr,” arXiv:2005.09394, 2020
H. Inaguma, M. Mimura, and T. Kawahara, · 2020
Closest in time.