Fetching the paper…
Reading the bibliography…
Recurrent neural transducer (RNN-T) is a promising end-to-end (E2E) model in automatic speech recognition (ASR).
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, et al., · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
R. Sennrich, B. Haddow, et al., · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, et al., · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, et al., · 2017
Earlier work this paper cites.
“A comparison of sequence-to-sequence models for speech recognition.,”
R. Prabhavalkar, K. Rao, et al., · 2017
Earlier work this paper cites.
“Exploring neural transducers for end-to-end speech recognition,”
E. Battenberg, J. Chen, et al., · 2017
Earlier work this paper cites.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping.,”
H. Sak, M. Shannon, et al., · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
H. Bu, J. Du, et al., · 2017
Cited alongside, same era.
“Deep-fsmn for large vocabulary continuous speech recognition,”
S. Zhang, M. Lei, et al., · 2018
Cited alongside, same era.
“Acoustic modeling with dfsmn-ctc and joint ctc-ce learning.,”
S. Zhang and M. Lei, · 2018
Cited alongside, same era.
“A comparison of end-to-end models for long-form speech recognition,”
C.-C. Chiu, W. Han, et al., · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T. N. Sainath, et al., · 2019
Cited alongside, same era.
“On the choice of modeling unit for sequence-to-sequence speech recognition,”
K. Irie, R. Prabhavalkar, et al., · 2019
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, et al., · 2019
Later among the works it cites.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
T. N. Sainath, Y. He, et al., · 2020
Closest in time.
W. Huang, W. Hu, Y. T. Yeung, and X. Chen, · 2020
Closest in time.
“Synchronous transformers for end-to-end speech recognition,” 2020
Z. Tian, J. Yi, et al., · 2020
Closest in time.
“Exploring pre-training with alignments for rnn transducer based end-to-end speech recognition,”
H. Hu, R. Zhao, et al., · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“An investigation into on-device personalization of end-to-end automatic speech recognition models,”
K. Chai Sim, P. Zadrazil, et al., · 2019
Cited alongside, same era.
“Transformers with convolutional context for asr,”
A. Mohamed, D. Okhonko, et al., · 2019
Cited alongside, same era.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Z. Dai, Z. Yang, et al., · 2019
Cited alongside, same era.
“Self-attention transducers for end-to-end speech recognition,”
Z. Tian, J. Yi, et al., · 2019
Cited alongside, same era.
W. Han, Z. Zhang, et al., · 2020
Closest in time.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, et al., · 2020
Closest in time.
“San-m: Memory equipped self-attention for end-to-end speech recognition,”
Z. Gao, S. Zhang, M. Lei, and I. McLoughlin, · 2020
Closest in time.
“Streaming chunk-aware multihead attention for online end-to-end speech recognition,” 2020
S. Zhang, Z. Gao, et al., · 2020
Closest in time.