Fetching the paper…
Reading the bibliography…
Monotonic chunkwise attention (MoChA) has been studied for the online streaming automatic speech recognition (ASR) based on a sequence-to-sequence framework.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “TED-LIUM: An automatic speech recognition dedicated corpus,” in
2012
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Earlier work this paper cites.
Y. Zhang, G. Chen, D. Yu, K. Yaco, S. Khudanpur, and J. Glass, “Highway long short-term memory RNNs for distant speech recognition,” in
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Earlier work this paper cites.
H. Sak, M. Shannon, K. Rao, and F. Beaufays, “Recurrent Neural Aligner: An encoder-decoder neural network model for sequence to sequence mapping,” in
2017
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in
2017
Earlier work this paper cites.
E. Battenberg, J. Chen, R. Child, A. Coates, Y. G. Y. Li, H. Liu, S. Satheesh, A. Sriram, and Z. Zhu, “Exploring neural transducers for end-to-end speech recognition,” in
2017
Cited alongside, same era.
C. Raffel, M.-T. Luong, P. J. Liu, R. J. Weiss, and D. Eck, “Online and linear-time attention by enforcing monotonic alignments,” in
2017
Cited alongside, same era.
S. Xue and Z. Yan, “Improving latency-controlled BLSTM acoustic models for online speech recognition,” in
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,”
2017
Cited alongside, same era.
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina
2018
Cited alongside, same era.
H. Miao, G. Cheng, P. Zhang, T. Li, and Y. Yan, “Online hybrid CTC/attention architecture for end-to-end speech recognition,” in
2019
Later among the works it cites.
R. Fan, P. Zhou, W. Chen, J. Jia, and G. Liu, “An online attention-based model for speech recognition,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in
2019
Later among the works it cites.
N. Moritz, T. Hori, and J. Le Roux, “Streaming end-to-end speech recognition with joint CTC-attention based models,” in
2019
Later among the works it cites.
Z. Tüske, K. Audhkhasi, and G. Saon, “Advancing sequence-to-sequence based speech recognition,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,” in
2018
Cited alongside, same era.
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, “An analysis of incorporating an external language model into a sequence-to-sequence model,” in
2018
Cited alongside, same era.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang
2019
Cited alongside, same era.
M. Huang, Y. Lu, L. Wang, Y. Qian, and K. Yu, “Exploring model units and training strategies for end-to-end speech recognition,” in
2019
Cited alongside, same era.
N. Moritz, T. Hori, and J. Le Roux, “Triggered attention for end-to-end speech recognition,” in
2019
Cited alongside, same era.
M. Li, M. Liu, and H. Masanori, “End-to-end speech recognition with adaptive computation steps,” in
2019
Cited alongside, same era.
K. Kim, K. Lee, D. Gowda, J. Park, S. Kim, S. Jin, Y.-Y. Lee, J. Yeo, D. Kim, S. Jung
2019
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga
2019
Later among the works it cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of Transformer and LSTM encoder decoder models for ASR,” in
2019
Later among the works it cites.
A. Garg, D. Gowda, A. Kumar, K. Kim, M. Kumar, and C. Kim, “Improved multi-stage training of online attention-based encoder-decoder models,” in
2019
Later among the works it cites.
L. Dong and B. Xu, “CIF: Continuous integrate-and-fire for end-to-end speech recognition,” in
2020
Closest in time.
H. Inaguma, Y. Gaur, L. Lu, J. Li, , and Y. Gong, “Minimum latency training strategies for streaming sequence-to-sequence ASR,” in
2020
Closest in time.
H. Miao, G. Cheng, P. Zhang, T. Li, and Y. Yan, “Online hybrid CTC/attention end-to-end automatic speech recognition architecture,”
2020
Closest in time.