Fetching the paper…
Reading the bibliography…
Recently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments.
“CSR-II (wsj1) complete,”
Linguistic Data Consortium, · 1994
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“CSR-I (wsj0) complete,”
John Garofalo, David Graff, Doug Paul, and David Pallett, · 2007
Earlier work this paper cites.
“Acoustic modeling using deep belief networks,”
Abdel-rahman Mohamed, George E Dahl, and Geoffrey Hinton, · 2012
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al., · 2012
Earlier work this paper cites.
“Adadelta: an adaptive learning rate method,”
Matthew D Zeiler, · 2012
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks,”
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio, · 2012
Cited alongside, same era.
“Hybrid speech recognition with deep bidirectional lstm,”
Alex Graves, Navdeep Jaitly, and Abdel-rahman Mohamed, · 2013
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Cited alongside, same era.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Cited alongside, same era.
“End-to-end continuous speech recognition using attention-based recurrent NN: First results,”
Jan Chorowski, Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Later among the works it cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Later among the works it cites.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals, · 2015
Later among the works it cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2015
Later among the works it cites.
“Chainer: a next-generation open source framework for deep learning,”
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton, · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc VV Le, · 2014
Cited alongside, same era.
“An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
Emmanuel Vincent, Shinji Watanabe, Aditya Arie Nugraha, Jon Barker, and Ricard Marxer,
Cited in the paper.
“Chainer,”
Preferred Networks,
Cited in the paper.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
Liang Lu, Xingxing Zhang, and Steve Renals, · 2016
Closest in time.
“On online attention-based speech recognition and joint mandarin character-pinyin training,”
William Chan and Ian Lane, · 2016
Closest in time.