Fetching the paper…
Reading the bibliography…
End-to-end models are gaining wider attention in the field of automatic speech recognition (ASR).
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
Y. Liu, P. Fung, Y. Yang, C. Cieri, S. Huang, and D. Graff, “Hkust/mts: A very large scale mandarin telephone speech corpus,” in
2006
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Earlier work this paper cites.
N. Jaitly, D. Sussillo, Q. V. Le, O. Vinyals, I. Sutskever, and S. Bengio, “A neural transducer,”
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
C. Raffel, M.-T. Luong, P. J. Liu, R. J. Weiss, and D. Eck, “Online and linear-time attention by enforcing monotonic alignments,” in
2017
Earlier work this paper cites.
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,”
2017
Cited alongside, same era.
J. Hou, S. Zhang, and L.-R. Dai, “Gaussian prediction based attention for online end-to-end speech recognition.” in
2017
Cited alongside, same era.
E. Battenberg, J. Chen, R. Child, A. Coates, Y. G. Y. Li, H. Liu, S. Satheesh, A. Sriram, and Z. Zhu, “Exploring neural transducers for end-to-end speech recognition,” in
2017
Cited alongside, same era.
H. Sak, M. Shannon, K. Rao, and F. Beaufays, “Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,” in
2017
Cited alongside, same era.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition.” in
2017
A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, “Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,” in
2019
Later among the works it cites.
N. Moritz, T. Hori, and J. Le Roux, “Triggered attention for end-to-end speech recognition,” in
2019
Later among the works it cites.
M. Li, M. Liu, and H. Masanori, “End-to-end speech recognition with adaptive computation steps,” in
2019
Later among the works it cites.
L. Dong, F. Wang, and B. Xu, “Self-attention aligner: A latency-control end-to-end model for asr using self-attention network and chunk-hopping,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” pp. 2613–2617, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Chorowski and N. Jaitly, “Towards better decoding and language model integration in sequence to sequence models,” pp. 523–527, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in
2018
Cited alongside, same era.
G. Pundak, T. N. Sainath, R. Prabhavalkar, A. Kannan, and D. Zhao, “Deep context: end-to-end contextual speech recognition,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
D. Duckworth, A. Neelakantan, B. Goodrich, L. Kaiser, and S. Bengio, “Parallel scheduled sampling,”
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Miao, G. Cheng, C. Gao, P. Zhang, and Y. Yan, “Transformer-based online ctc/attention end-to-end speech recognition architecture,” in
2020
Closest in time.
L. Dong and B. Xu, “Cif: Continuous integrate-and-fire for end-to-end speech recognition,” in
2020
Closest in time.