Fetching the paper…
Reading the bibliography…
In Automatic Speech Recognition it is still challenging to learn useful intermediate representations when using high-level (or abstract) target units such as words.
“An introduction to hidden markov models,”
Lawrence Rabiner and B Juang, · 1986
Earlier work this paper cites.
“A new algorithm for data compression,”
Philip Gage, · 1994
Earlier work this paper cites.
“Multitask learning,”
Rich Caruana, · 1997
Earlier work this paper cites.
“A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),”
Jonathan G Fiscus, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky and James Martin, · 2000
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jurgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence labelling in structured domains with hierarchical recurrent neural networks,”
Santiago Fernández, Alex Graves, and Jürgen Schmidhuber, · 2007
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Subword language modeling with neural networks,”
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and Jan Cernocky, · 2012
Earlier work this paper cites.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Cited alongside, same era.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Latent sequence decompositions,”
Quoc V. Le Navdeep Jaitly William Chan, Yu Zhang, · 2016
Cited alongside, same era.
“Deep multi-task learning with low level tasks supervised at lower layers,”
Anders Søgaard and Yoav Goldberg, · 2016
Cited alongside, same era.
“Character-level incremental speech recognition with recurrent neural networks,”
Kyuyeon Hwang and Wonyong Sung, · 2016
“Gram-CTC: Automatic unit selection and target decomposition for sequence labelling,”
Hairong Liu, Zhenyao Zhu, Xiangang Li, and Sanjeev Satheesh, · 2017
Later among the works it cites.
“Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition,”
Takaaki Hori, Shinji Watanabe, and John R Hershey, · 2017
Later among the works it cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Later among the works it cites.
“Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition,”
Shubham Toshniwal, Hao Tang, Liang Lu, and Karen Livescu, · 2017
Later among the works it cites.
“Comparison of decoding strategies for CTC acoustic models,”
Thomas Zenkel, Ramon Sanabria, Florian Metze, Jan Niehues, Matthias Sperber, Sebastian Stüker, and Alex Waibel, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Katya Gonina, et al., · 2017
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition,”
Hagen Soltau, Hank Liao, and Hasim Sak, · 2017
Cited alongside, same era.
“Direct acoustics-to-word models for english conversational speech recognition,”
Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, and David Nahamoo, · 2017
Cited alongside, same era.
“Subword and crossword units for CTC acoustic models,”
Thomas Zenkel, Ramon Sanabria, Florian Metze, and Alex Waibel, · 2017
Cited alongside, same era.
“Building competitive direct acoustics-to-word models for english conversational speech recognition,”
Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Michael Picheny, · 2018
Closest in time.
“Sequence-based multi-lingual low resource speech recognition,”
Siddharth Dalmia, Ramon Sanabria, Florian Metze, and Alan W Black, · 2018
Closest in time.
“Hierarchical multitask learning for CTC-based speech recognition,”
Kalpesh Krishna, Shubham Toshniwal, and Karen Livescu, · 2018
Closest in time.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Closest in time.