Fetching the paper…
Reading the bibliography…
This paper proposes and evaluates the hybrid autoregressive transducer (HAT) model, a time-synchronous encoderdecoder model that preserves the modularity of conventional automatic speech recognition systems.
“Continuous speech recognition by statistical methods,”
Frederick Jelinek, · 1976
Earlier work this paper cites.
“A maximum likelihood approach to continuous speech recognition,”
Lalit R Bahl, Frederick Jelinek, and Robert L Mercer, · 1983
Earlier work this paper cites.
“Maximum mutual information estimation of hidden markov model parameters for speech recognition,”
Lalit R Bahl, Peter F Brown, et al., · 1986
Earlier work this paper cites.
“Memoir on the probability of the causes of events,”
Pierre Simon Laplace, · 1986
Earlier work this paper cites.
“Continuous speech recognition using multilayer perceptrons with hidden markov models,”
Nelson Morgan and Herve Bourlard, · 1990
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Discriminative training on language model,”
Zheng Chen, Kai-Fu Lee, and Ming-jing Li, · 2000
Earlier work this paper cites.
“Weighted finite-state transducers in speech recognition,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2002
Earlier work this paper cites.
“Discriminative training of language models for speech recognition,”
Hong-Kwang Jeff Kuo, Eric Fosler-Lussier, et al., · 2002
Earlier work this paper cites.
Discriminative training for large vocabulary speech recognition
Daniel Povey, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, et al., · 2006
Earlier work this paper cites.
“Speech recognition with weighted finite-state transducers,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2008
Earlier work this paper cites.
“Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,”
Brian Kingsbury, · 2009
Earlier work this paper cites.
“From speech to letters-using a novel neural network architecture for grapheme based asr,”
Florian Eyben, Martin Wöllmer, et al., · 2009
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Joint language and translation modeling with recurrent neural networks,”
Michael Auli, Michel Galley, et al., · 2013
Cited alongside, same era.
“Recurrent continuous translation models,”
Nal Kalchbrenner and Phil Blunsom, · 2013
Cited alongside, same era.
“Learning phrase representations using rnn encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, et al., · 2014
Cited alongside, same era.
“On the properties of neural machine translation: Encoder-decoder approaches,”
Kyunghyun Cho, Bart Van Merriënboer, et al., · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2016
Later among the works it cites.
“Comparison of decoding strategies for ctc acoustic models,”
Thomas Zenkel, Ramon Sanabria, et al., · 2017
Later among the works it cites.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, et al., · 2017
Later among the works it cites.
“Cold fusion: Training seq2seq models together with language models,”
Anuroop Sriram, Heewoo Jun, et al., · 2017
Later among the works it cites.
“End-to-end training of acoustic models for large vocabulary continuous speech recognition with tensorflow,”
Ehsan Variani, Tom Bagby, et al., · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Awni Hannun, Carl Case, et al., · 2014
Cited alongside, same era.
“Rapid vocabulary addition to context-dependent decoder graphs,”
Cyril Allauzen and Michael Riley, · 2015
Cited alongside, same era.
“Composition-based on-the-fly rescoring for salient n-gram biasing,”
Keith Hall, Eunjoon Cho, et al., · 2015
Cited alongside, same era.
“On using monolingual corpora in neural machine translation,”
Caglar Gulcehre, Orhan Firat, et al., · 2015
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Haşim Sak, Andrew Senior, et al., · 2015
Cited alongside, same era.
“Sequence level training with recurrent neural networks,”
Marc’Aurelio Ranzato, Sumit Chopra, et al., · 2015
Cited alongside, same era.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Cited alongside, same era.
Later among the works it cites.
“Efficient implementation of the room simulator for training deep neural network acoustic models,”
Chanwoo Kim, Ehsan Variani, et al., · 2017
Later among the works it cites.
“Effectively building tera scale maxent language models incorporating non-linguistic signals,”
Fadi Biadsy, Mohammadreza Ghodsi, and Diamantino Caseiro, · 2017
Later among the works it cites.
“End-to-end speech recognition with word-based rnn language models,”
Takaaki Hori, Jaejin Cho, and Shinji Watanabe, · 2018
Later among the works it cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, et al., · 2018
Later among the works it cites.
“No need for a lexicon? evaluating the value of the pronunciation lexica in end-to-end models,”
Tara N Sainath, Rohit Prabhavalkar, et al., · 2018
Later among the works it cites.
“A comparison of modeling units in sequence-to-sequence speech recognition with the transformer on mandarin chinese,”
Shiyu Zhou, Linhao Dong, et al., · 2018
Later among the works it cites.
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
Changhao Shan, Chao Weng, et al., · 2019
Later among the works it cites.
“Model unit exploration for sequence-to-sequence speech recognition,”
Kazuki Irie, Rohit Prabhavalkar, et al., · 2019
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, et al., · 2019
Later among the works it cites.