Fetching the paper…
Reading the bibliography…
In automatic speech recognition (ASR) rescoring, the hypothesis with the fewest errors should be selected from the n-best list using a language model (LM).
“Flexible speech understanding based on combined key-phrase detection and verification,”
Tatsuya Kawahara, Chin-Hui Lee, and B. Juang, · 1998
Earlier work this paper cites.
“Evaluation of word confidence for speech recognition systems,”
M. Siu and H. Gish, · 1999
Earlier work this paper cites.
“Confidence measures for large vocabulary continuous speech recognition,”
F. Wessel, R. Schluter, K. Macherey, and H. Ney, · 2001
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks,”
Anthony Rousseau, Paul Deléglise, and Yannick Estève, · 2014
Earlier work this paper cites.
“Discriminative method for recurrent neural network language models,”
Yuuki Tachioka and Shinji Watanabe, · 2015
Earlier work this paper cites.
“Bidirectional recurrent neural network language models for automatic speech recognition,”
Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, and Stanley Chen, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Minimum word error training of long short-term memory recurrent neural network language models for speech recognition,”
Takaaki Hori, Chiori Hori, Shinji Watanabe, and John R. Hershey, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Earlier work this paper cites.
“Investigating bidirectional recurrent neural network language models for speech recognition,”
X. Chen, A. Ragni, X. Liu, and Mark J.F. Gales, · 2017
Earlier work this paper cites.
“On calibration of modern neural networks,”
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger, · 2017
Earlier work this paper cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, ZhiJeng Chen, and Rohit Prabhavalkar, · 2018
Cited alongside, same era.
“Large margin neural language model,”
Jiaji Huang, Yi Li, Wei Ping, and Liang Huang, · 2018
Cited alongside, same era.
“Confidence estimation and deletion prediction using bidirectional recurrent neural networks,”
Anton Ragni, Qiujia Li, Mark Gales, and Yu Wang, · 2018
Cited alongside, same era.
“Deep contextualized word representations,”
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Cited alongside, same era.
“Language modeling with deep transformers,”
“Parallel rescoring with transformer for streaming on-device speech recognition,”
Wei Li, James Qin, Chung-Cheng Chiu, Ruoming Pang, and Yanzhang He, · 2020
Later among the works it cites.
“ELECTRA: Pre-training text encoders as discriminators rather than generators,”
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning, · 2020
Later among the works it cites.
“Confidence measures in encoder-decoder models for speech recognition,”
Alejandro Woodward, Clara Bonnín, Issey Masuda, David Varas, Elisenda Bou-Balust, and Juan Carlos Riveiro, · 2020
Later among the works it cites.
“Pre-training transformers as energy-based cloze models,”
Kevin Clark, Minh-Thang Luong, Quoc Le, and Christopher D. Manning, · 2020
Later among the works it cites.
“Audio-attention discriminative language model for asr rescoring,”
Ankur Gandhe and Ariya Rastrow, · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Effective sentence scoring method using BERT for speech recognition,”
Joonbo Shin, Yoonhyung Lee, and Kyomin Jung, · 2019
Cited alongside, same era.
“Improving ASR confidence scores for alexa using acoustic and hypothesis embeddings,”
Prakhar Swarup, Roland Maas, Sri Garimella, Sri Harish Mallidi, and Björn Hoffmeister, · 2019
Cited alongside, same era.
“BERT has a mouth, and it must speak: BERT as a Markov random field language model,”
Alex Wang and Kyunghyun Cho, · 2019
Cited alongside, same era.
“Mask-predict: Parallel decoding of conditional masked language models,”
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
Daniel S. Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V. Le, · 2020
Later among the works it cites.
“Phoneme-to-grapheme conversion based large-scale pre-training for end-to-end automatic speech recognition,”
Ryo Masumura, Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka, and Shota Orihashi, · 2020
Later among the works it cites.
“Transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush, · 2020
Later among the works it cites.
“BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,”
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer, · 2020
Later among the works it cites.
“Confidence estimation for attention-based sequence-to-sequence models for speech recognition,”
Qiujia Li, David Qiu, Yu Zhang, Bo Li, Yanzhang He, Philip C. Woodland, Liangliang Cao, and Trevor Strohman, · 2021
Closest in time.
“Blstm-based confidence estimation for end-to-end speech recognition,”
Atsunori Ogawa, Naohiro Tawara, Takatomo Kano, and Marc Delcroix, · 2021
Closest in time.
“An evaluation of word-level confidence estimation for end-to-end automatic speech recognition,”
Dan Oneaţă, Alexandru Caranica, Adriana Stan, and Horia Cucu, · 2021
Closest in time.
“Large margin training improves language models for ASR,”
Jilin Wang, Jiaji Huang, and Kenneth Ward Church, · 2021
Closest in time.
“Learning word-level confidence for subword end-to-end ASR,”
David Qiu, Qiujia Li, Yanzhang He, Yu Zhang, Bo Li, Liangliang Cao, Rohit Prabhavalkar, Deepti Bhatia, Wei Li, Ke Hu, Tara N. Sainath, and Ian McGraw, · 2021
Closest in time.
“A general multi-task learning framework to leverage text data for speech to text tasks,”
Yun Tang, Juan Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel, · 2021
Closest in time.
“Cascade rnn-transducer: Syllable based streaming on-device mandarin speech recognition with a syllable-to-character converter,”
Xiong Wang, Zhuoyuan Yao, Xian Shi, and Lei Xie, · 2021
Closest in time.
“Making punctuation restoration robust and fast with multi-task learning and knowledge distillation,”
Michael Hentschel, Emiru Tsunoo, and Takao Okuda, · 2021
Closest in time.