Fetching the paper…
Reading the bibliography…
One challenging problem of robust automatic speech recognition (ASR) is how to measure the goodness of a speech enhancement algorithm (SEA) without calculating the word error rate (WER) due to the high costs of manual transcriptions, language modeling and decoding process.
“Acoustic confidence measures for segmenting broadcast news,”
Jon Barker, Gethin Williams, and Steve Renals, · 1998
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
Correlation: Parametric and nonparametric measures
Peter Y Chen and Paula M Popovich, · 2002
Earlier work this paper cites.
“New entropy based combination rules in HMM/ANN multi-stream ASR,”
Hemant Misra, Hervé Bourlard, and Vivek Tyagi, · 2003
Earlier work this paper cites.
“Noise spectrum estimation in adverse environments: Improved minima controlled recursive averaging,”
Israel Cohen, · 2003
Earlier work this paper cites.
“Performance estimation of speech recognition system under noise conditions using objective quality measures and artificial voice,”
Takeshi Yamada, Masakazu Kumakura, and Nobuhiko Kitawaki, · 2006
Earlier work this paper cites.
“Objective quality evaluation in blind source separation for speech recognition in a real room,”
Leandro Di Persia, Masuzo Yanagida, Hugo Leonardo Rufiner, and Diego Milone, · 2007
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, · 2011
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Performance estimation of noisy speech recognition using spectral distortion and SNR of noise-reduced speech,”
Guo Ling, Takeshi Yamada, Shoji Makino, and Nobuhiko Kitawaki, · 2013
Cited alongside, same era.
“Estimation of speech recognition performance in noisy and reverberant environments using PESQ score and acoustic parameters,”
Takahiro Fukumori, Masato Nakayama, Takanobu Nishiura, and Yoichi Yamashita, · 2013
Cited alongside, same era.
“Ideal ratio mask estimation using deep neural networks for robust speech recognition,”
Arun Narayanan and DeLiang Wang, · 2013
Cited alongside, same era.
“An overview of noise-robust automatic speech recognition,”
Jinyu Li, Li Deng, Yifan Gong, and Reinhold Haeb-Umbach, · 2014
Cited alongside, same era.
“Environmentally robust asr front-end for deep neural network acoustic models,”
Takuya Yoshioka and Mark JF Gales, · 2015
Cited alongside, same era.
“Estimating speech recognition accuracy based on error type classification,”
Atsunori Ogawa, Takaaki Hori, and Atsushi Nakamura, · 2016
Later among the works it cites.
“Deep convolutional neural networks with layer-wise context expansion and attention.,”
Dong Yu, Wayne Xiong, Jasha Droppo, Andreas Stolcke, Guoli Ye, Jinyu Li, and Geoffrey Zweig, · 2016
Later among the works it cites.
“Speech enhancement for robust automatic speech recognition: Evaluation using a baseline system and instrumental measures,”
Alastair H Moore, P Peso Parada, and Patrick A Naylor, · 2017
Later among the works it cites.
“Recent progresses in deep learning based acoustic models,”
Dong Yu and Jinyu Li, · 2017
Later among the works it cites.
“An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
Emmanuel Vincent, Shinji Watanabe, Aditya Arie Nugraha, Jon Barker, and Ricard Marxer, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speech enhancement and noise-robust automatic speech recognition,”
DA Thomsen and Carina E Andersen, · 2015
Cited alongside, same era.
“Combining spectral feature mapping and multi-channel model-based source separation for noise-robust automatic speech recognition,”
Deblin Bagchi, Michael I Mandel, Zhongqiu Wang, Yanzhang He, Andrew Plummer, and Eric Fosler-Lussier, · 2015
Cited alongside, same era.
“Performance estimation of noisy speech recognition using spectral distortion and recognition task complexity,”
Ling Guo, Takeshi Yamada, Shigeki Miyabe, Shoji Makino, and Nobuhiko Kitawaki, · 2016
Cited alongside, same era.
Szu-Jui Chen, Aswin Shanmugam Subramanian, Hainan Xu, and Shinji Watanabe, · 2018
Closest in time.
“Error modeling via asymmetric laplace distribution for deep neural network based single-channel speech enhancement,”
Li Chai, Jun Du, and Chin-Hui Lee, · 2018
Closest in time.
“End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,”
Szu-Wei Fu, Tao-Wei Wang, Yu Tsao, Xugang Lu, and Hisashi Kawai, · 2018
Closest in time.