Fetching the paper…
Reading the bibliography…
Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
Karel Veselỳ, Arnab Ghoshal, Lukás Burget, and Daniel Povey, · 2013
Earlier work this paper cites.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
Hagen Soltau, Hank Liao, and Hasim Sak, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Sequence student-teacher training of deep neural networks,”
JHM Wong and MJF Gales, · 2016
Earlier work this paper cites.
“Simplifying long short-term memory acoustic models for fast training and decoding,”
Yajie Miao, Jinyu Li, Yongqiang Wang, Shi-Xiong Zhang, and Yifan Gong, · 2016
Earlier work this paper cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Advancing acoustic-to-word ctc model,”
Jinyu Li, Guoli Ye, Amit Das, Rui Zhao, and Yifan Gong, · 2018
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, et al., · 2018
Earlier work this paper cites.
“A comparison of techniques for language model integration in encoder-decoder speech recognition,”
Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu, Yonghui Wu, Tara N Sainath, and Karen Livescu, · 2018
Earlier work this paper cites.
“Deep context: end-to-end contextual speech recognition,”
Golan Pundak, Tara N Sainath, Rohit Prabhavalkar, Anjuli Kannan, and Ding Zhao, · 2018
Earlier work this paper cites.
“Subword regularization: Improving neural network translation models with multiple subword candidates,”
Taku Kudo, · 2018
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, et al., · 2019
Cited alongside, same era.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
Erik McDermott, Hasim Sak, and Ehsan Variani, · 2019
Cited alongside, same era.
“Shallow-fusion end-to-end contextual biasing.,”
Ding Zhao, Tara N Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang, · 2019
Cited alongside, same era.
“Integrating source-channel and attention-based sequence-to-sequence models for speech recognition,”
Qiujia Li, Chao Zhang, and Philip C Woodland, · 2019
Cited alongside, same era.
“Layer trajectory blstm.,”
“High-accuracy and low-latency speech recognition with two-head contextual layer trajectory LSTM model,”
Jinyu Li, Rui Zhao, Eric Sun, Jeremy HM Wong, Amit Das, Zhong Meng, and Yifan Gong, · 2020
Later among the works it cites.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Closest in time.
“An efficient streaming non-recurrent on-device end-to-end model with improvements to rare-word modeling,”
Tara N. Sainath, Yanzhang He, Arun Narayanan, et al., · 2021
Closest in time.
“Recent advances in end-to-end automatic speech recognition,”
Jinyu Li, · 2021
Closest in time.
“Internal language model estimation for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, Rui Zhao, Jinyu Li, and Yifan Gong, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eric Sun, Jinyu Li, and Yifan Gong, · 2019
Cited alongside, same era.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Cited alongside, same era.
“On the comparison of popular end-to-end models for large scale speech recognition,”
Jinyu Li, Yu Wu, Yashesh Gaur, Chengyi Wang, Rui Zhao, and Shujie Liu, · 2020
Cited alongside, same era.
“Combination of end-to-end and hybrid models for speech recognition.,”
Jeremy HM Wong, Yashesh Gaur, Rui Zhao, Liang Lu, Eric Sun, Jinyu Li, and Yifan Gong, · 2020
Cited alongside, same era.
“Deliberation model based two-pass end-to-end speech recognition,”
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, et al., · 2020
Cited alongside, same era.
Closest in time.
Xiaoqiang Wang, Yanqing Liu, Sheng Zhao, and Jinyu Li, · 2021
Closest in time.
“A better and faster end-to-end model for streaming asr,”
Bo Li, Anmol Gulati, Jiahui Yu, et al., · 2021
Closest in time.
“Reducing streaming asr model delay with self alignment,”
Jaeyoung Kim, Han Lu, Anshuman Tripathi, Qian Zhang, and Hasim Sak, · 2021
Closest in time.
“Combining frame-synchronous and label-synchronous systems for speech recognition,”
Qiujia Li, Chao Zhang, and Philip C Woodland, · 2021
Closest in time.
“Transformer based deliberation for two-pass speech recognition,”
Ke Hu, Ruoming Pang, Tara N Sainath, and Trevor Strohman, · 2021
Closest in time.
“Cascaded encoders for unifying streaming and non-streaming asr,”
Arun Narayanan, Tara N Sainath, Ruoming Pang, Jiahui Yu, Chung-Cheng Chiu, Rohit Prabhavalkar, Ehsan Variani, and Trevor Strohman, · 2021
Closest in time.