Fetching the paper…
Reading the bibliography…
In this work, we perform an empirical comparison among the CTC, RNN-Transducer, and attention-based Seq2Seq models for end-to-end speech recognition.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
The Kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, K. Veselý, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, and G. Stemmer · 2011
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G.E. Hinton, L. Deng, D. Yu, G.E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Dirt cheap web-scale parallel text from the common crawl
Jason R Smith, Herve Saint-Amand, Magdalena Plamada, Philipp Koehn, Chris Callison-Burch, and Adam Lopez · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2014
Earlier work this paper cites.
First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs
Awni Y. Hannun, Andrew L. Maas, Daniel Jurafsky, and Andrew Y. Ng · 2014
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2015
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2015
Earlier work this paper cites.
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2015
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals · 2015
Cited alongside, same era.
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitry Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding
Yajie Miao, Mohammad Gowayyed, and Florian Metze · 2015
Cited alongside, same era.
Acoustic modelling with cd-ctc-smbr lstm rnns
Andrew W. Senior, Hasim Sak, Felix de Chaumont Quitry, Tara N. Sainath, and Kanishka Rao · 2015
Cited alongside, same era.
Dense prediction on sequences with time-dilated convolutions for speech recognition
Tom Sercu and Vaibhava Goel · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Later among the works it cites.
Achieving human parity in conversational speech recognition
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig · 2016
Later among the works it cites.
Advances in all-neural speech recognition
Geoffery Zweig, Ghengzhu Yu, Jasha Droppo, and Andreas Stolcke · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sequence to sequence transduction with hard monotonic attention
Roee Aharoni and Yoav Goldberg · 2016
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Cited alongside, same era.
Segmental recurrent neural networks for end-to-end speech recognition
Liang Lu, Lingpeng Kong, Chris Dyer, Noah A. Smith, and Steve Renals · 2016
Cited alongside, same era.
Purely sequence-trained neural networks for asr based on lattice-free mmi
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur · 2016
Cited alongside, same era.
Eric Battenberg, Rewon Child, Adam Coates, Christopher Fougner, Yashesh Gaur, Jiaji Huang, Heewoo Jun, Ajay Kannan, Markus Kliegl, Atul Kumar, et al · 2017
Closest in time.
An online sequence-to-sequence model for noisy speech recognition
Chung-Cheng Chiu, Dieterich Lawson, Yuping Luo, George Tucker, Kevin Swersky, Ilya Sutskever, and Navdeep Jaitly · 2017
Closest in time.
Gram-ctc: Automatic unit selection and target decomposition for sequence labelling
Hairong Liu, Zhenyao Zhu, Xiangang Li, and Sanjeev Satheesh · 2017
Closest in time.
Online and linear-time attention by enforcing monotonic alignments
Colin Raffel, Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck · 2017
Closest in time.
English conversational telephone speech recognition by humans and machines
George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, et al · 2017
Closest in time.