Fetching the paper…
Reading the bibliography…
Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks.
Discriminative training for large vocabulary speech recognition
Daniel Povey, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,”
Brian Kingsbury, · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
Karel Veselỳ, Arnab Ghoshal, Lukás Burget, and Daniel Povey, · 2013
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Earlier work this paper cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“End-to-end speech recognition using lattice-free mmi,”
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Minimum word error rate training for attention-based sequence-to-sequence models,”
Rohit Prabhavalkar, Tara N. Sainath, Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-Cheng Chiu, and Anjuli Kannan, · 2018
Cited alongside, same era.
“Improving attention based sequence-to-sequence models for end-to-end english conversational speech recognition.,”
Chao Weng, Jia Cui, Guangsen Wang, Jun Wang, Chengzhu Yu, Dan Su, and Dong Yu, · 2018
Cited alongside, same era.
“End-to-end speech recognition with word-based rnn language models,”
Takaaki Hori, Jaejin Cho, and Shinji Watanabe, · 2018
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Later among the works it cites.
“Efficient minimum word error rate training of rnn-transducer for end-to-end speech recognition,”
Jinxi Guo, Gautam Tiwari, Jasha Droppo, Maarten Van Segbroeck, Che-Wei Huang, and Stolcke, · 2020
Later among the works it cites.
“Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,”
Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang, Zhong Meng, Zhuo Chen, Tianyan Zhou, and Takuya Yoshioka, · 2020
Later among the works it cites.
“Alignment-length synchronous decoding for rnn transducer,”
George Saon, Zoltán Tüske, and Kartik Audhkhasi, · 2020
Later among the works it cites.
“Multi-head monotonic chunkwise attention for online speech recognition,”
Baiji Liu, Songjun Cao, Sining Sun, Weibin Zhang, and Long Ma, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Cited alongside, same era.
“An overview of end-to-end automatic speech recognition,”
Dong Wang, Xiaodong Wang, and Shaohe Lv, · 2019
Cited alongside, same era.
“Minimum bayes risk training of rnn-transducer for end-to-end speech recognition,”
Chao Weng, Chengzhu Yu, Jia Cui, Chunlei Zhang, and Dong Yu, · 2019
Cited alongside, same era.
Later among the works it cites.
“Minimum bayes risk training for end-to-end speaker-attributed asr,”
Naoyuki Kanda, Zhong Meng, Liang Lu, Yashesh Gaur, Xiaofei Wang, Zhuo Chen, and Takuya Yoshioka, · 2021
Closest in time.
“Non-autoregressive transformer asr with ctc-enhanced decoder input,”
Xingchen Song, Zhiyong Wu, Yiheng Huang, Chao Weng, Dan Su, and Helen Meng, · 2021
Closest in time.