Fetching the paper…
Reading the bibliography…
Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria.
“Maximum mutual information estimation of hidden markov model parameters for speech recognition,”
L. Bahl, P. Brown, P. de Souza, and R. Mercer, · 1986
Earlier work this paper cites.
“A time-delay neural network architecture for isolated word recognition,”
Kevin J Lang, Alex H Waibel, and Geoffrey E Hinton, · 1990
Earlier work this paper cites.
“Tree-based state tying for high accuracy acoustic modelling,”
S. J. Young, J. J. Odell, and P. C. Woodland, · 1994
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, and Faustino Gomez, · 2006
Earlier work this paper cites.
“Boosted mmi for model and feature-space discriminative training,”
Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Karthik Visweswariah, · 2008
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
George E. Dahl, Dong Yu, Li Deng, and Alex Acero, · 2012
Earlier work this paper cites.
“Discriminative feature-space transforms using deep neural networks,”
George Saon and Brian Kingsbury, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Semi-supervised training of deep neural networks,”
K. Vesely, M. Hannemann, and L. Burget, · 2013
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Learning acoustic frame labeling for speech recognition with recurrent neural networks,”
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan İrsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk, · 2015
Earlier work this paper cites.
“Pronunciation and silence probability modeling for asr,”
Guoguo Chen, Hainan Xu, Minhua Wu, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Semi-supervised maximum mutual information training of deep neural network acoustic models,”
Vimal Manohar, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Purely sequence-trained neural networks for asr based on lattice-free mmi,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philémon Brakel, and Yoshua Bengio, · 2016
Cited alongside, same era.
“A comparison of lattice-free discriminative training criteria for purely sequence-trained neural network acoustic models,”
Chao Weng and Dong Yu, · 2019
Later among the works it cites.
“Promising accurate prefix boosting for sequence-to-sequence asr,”
Murali Karthick Baskar, Lukáš Burget, Shinji Watanabe, Martin Karafiát, Takaaki Hori, and Jan Honza Černockỳ, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.
Ashish Arora, Chun Chieh Chang, Babak Rekabdar, Bagher BabaAli, Daniel Povey, David Etter, Desh Raj, Hossein Hadian, Jan Trmal, Paola Garcia, et al., · 2019
Later among the works it cites.
“A New Training Pipeline for an Improved Neural Transducer,”
Albert Zeyer, André Merboldt, Ralf Schlüter, and Hermann Ney, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Flat-start single-stage discriminatively trained hmm-based models for asr,”
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
Cited alongside, same era.
“Semi-supervised training of acoustic models using Lattice-Free MMI,”
V. Manohar, H. Hadian, D. Povey, and S. Khudanpur, · 2018
Cited alongside, same era.
“Sequence discriminative training for deep learning based acoustic keyword spotting,”
Zhehuai Chen, Yanmin Qian, and Kai Yu, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
Later among the works it cites.
“Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces,”
Frank Zhang, Yongqiang Wang, Xiaohui Zhang, Chunxi Liu, Yatharth Saraf, and Geoffrey Zweig, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung Cheng Chiu, and Others, · 2020
Later among the works it cites.
“Benchmarking lf-mmi, ctc and rnn-t criteria for streaming asr,”
Xiaohui Zhang, Frank Zhang, Chunxi Liu, Kjell Schubert, Julian Chan, Pradyot Prakash, Jun Liu, Ching-Feng Yeh, Fuchun Peng, Yatharth Saraf, and Geoffrey Zweig, · 2021
Closest in time.
“Improving rnn transducer based asr with auxiliary tasks,”
Chunxi Liu, Frank Zhang, Duc Le, Suyoun Kim, Yatharth Saraf, and Geoffrey Zweig, · 2021
Closest in time.
“Alignment restricted streaming recurrent neural network transducer,”
Jay Mahadeokar, Yuan Shangguan, Duc Le, Gil Keren, Hang Su, Thong Le, Ching-Feng Yeh, Christian Fuegen, and Michael L Seltzer, · 2021
Closest in time.
Duc Le, Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fuegen, Ozlem Kalinli, Yatharth Saraf, and Michael L. Seltzer, · 2021
Closest in time.
“Why does ctc result in peaky behavior?,”
Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2021
Closest in time.
“Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition,”
Yangyang Shi, Yongqiang Wang, Chunyang Wu, Ching-Feng Yeh, and Others, · 2021
Closest in time.