Fetching the paper…
Reading the bibliography…
The attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network.
“Hkust/mts: A very large scale mandarin telephone speech corpus.,”
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff, · 2006
Earlier work this paper cites.
“Adadelta: An adaptive learning rate method,”
Matthew D. Zeiler, · 2012
Earlier work this paper cites.
“Understanding the exploding gradient problem,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski et al., · 2015
Earlier work this paper cites.
“Deep speech 2 : End-to-end speech recognition in english and mandarin,”
Dario Amodei et al., · 2015
Earlier work this paper cites.
“Semi-supervised sequence learning,”
Andrew M Dai and Quoc V Le, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau et al., · 2016
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“A comparison of sequence-to-sequence models for speech recognition,”
Rohit Prabhavalkar et al., · 2017
Cited alongside, same era.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Cited alongside, same era.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Knowing when to look: Adaptive attention via a visual sentinel for image captioning,”
“Semi-supervised training for improving data efficiency in end-to-end speech synthesis,”
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and R. J. Skerry-Ryan, · 2018
Later among the works it cites.
“End-to-end speech recognition with word-based rnn language models,”
Takaaki Hori, Jaejin Cho, and Shinji Watanabe, · 2018
Later among the works it cites.
“Multi-modal data augmentation for end-to-end asr,”
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner, and Shinji Watanabe, · 2018
Later among the works it cites.
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe et al., · 2018
Later among the works it cites.
“Catastrophic forgetting: still a problem for dnns,”
Benedikt Pfülb, Alexander Gepperth, S. Abdullah, and A. Kilian, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lu, C. Xiong, D. Parikh, and R. Socher, · 2017
Cited alongside, same era.
“Unsupervised pretraining for sequence to sequence learning,”
Prajit Ramachandran, Peter J. Liu, and Quoc V. Le, · 2017
Cited alongside, same era.
“Advances in joint ctc-attention based end-to-end speech recognition with a deep cnn encoder and rnn-lm,”
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan, · 2017
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu et al., · 2018
Cited alongside, same era.
“A spelling correction model for end-to-end speech recognition,”
Jinxi Guo, Tara N. Sainath, and Ron J. Weiss, · 2019
Closest in time.
“MASS: masked sequence to sequence pre-training for language generation,”
Kaitao Song et al., · 2019
Closest in time.
“Building the singapore english national speech corpus,”
Jia Xin Koh et al., · 2019
Closest in time.