Fetching the paper…
Reading the bibliography…
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality.
“Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,”
Y. Ephraim and D. Malah, · 1984
Earlier work this paper cites.
“Estimation of modal decay parameters from noisy response measurements,”
Poju Antsalo et al., · 2001
Earlier work this paper cites.
Feb 2001
“ITU-T recommendation P.862: Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” · 2001
Earlier work this paper cites.
“The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings,”
Joachim Thiemann, Nobutaka Ito, and Emmanuel Vincent, · 2013
Earlier work this paper cites.
“Perceptual objective listening quality assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement part II-perceptual model,”
John Beerends et al., · 2013
Earlier work this paper cites.
“CREMA-D: Crowd-sourced emotional multimodal actors dataset,”
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma, · 2014
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“THCHS-30 : A free chinese speech corpus,” 2015
Zhiyong Zhang Dong Wang, Xuewei Zhang, · 2015
Earlier work this paper cites.
“Speaker recognition by machines and humans: A tutorial review,”
John HL Hansen and Taufiq Hasan, · 2015
Earlier work this paper cites.
“An individualized super-gaussian single microphone speech enhancement for hearing aid users with smartphone as an assistive device,”
C. Karadagur Ananda Reddy, N. Shankar, G. Shreedhar Bhat, R. Charan, and I. Panahi, · 2017
Earlier work this paper cites.
“Raw waveform-based speech enhancement by fully convolutional networks,”
S. Fu, Y. Tsao, X. Lu, and H. Kawai, · 2017
Cited alongside, same era.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Cited alongside, same era.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Cited alongside, same era.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Cited alongside, same era.
“Vocalset: A singing voice dataset.,”
Julia Wilkins, Prem Seetharaman, Alison Wahl, and Bryan Pardo, · 2018
Cited alongside, same era.
“Voxceleb2: Deep speaker recognition,”
Yuichiro Koyama, Tyler Vuong, Stefan Uhlich, and Bhiksha Raj, · 2020
Closest in time.
“A perceptually-motivated approach for low-complexity, real-time enhancement of fullband speech,”
Jean-Marc Valin et al., · 2020
Closest in time.
Umut Isik et al., · 2020
Closest in time.
“The INTERSPEECH 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,”
Chandan KA Reddy et al., · 2020
Closest in time.
“An open source implementation of ITU-T recommendation P.808 with validation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Cited alongside, same era.
“A scalable noisy speech dataset and online subjective test framework,”
Chandan KA Reddy et al., · 2019
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al., · 2019
Cited alongside, same era.
“Non-intrusive speech quality assessment using neural networks,”
A. R. Avila, H. Gamper, C. Reddy, R. Cutler, I. Tashev, and J. Gehrke, · 2019
Cited alongside, same era.
“Phase-aware single-stage speech denoising and dereverberation with U-net,”
Hyeong-Seok Choi, Hoon Heo, Jie Hwan Lee, and Kyogu Lee, · 2020
Cited alongside, same era.
Babak Naderi and Ross Cutler, · 2020
Closest in time.
[Online; accessed 2020-09-01]
“The Spoken Wikipedia Corpora,” https://nats.gitlab.io/swc/ , · 2020
Closest in time.
[Online; accessed 2020-09-01]
“Telecooperation German Corpus for Kinect,” http://www.repository.voxforge1.org/downloads/de/german-speechdata-TUDa-2015.tar.gz , · 2020
Closest in time.
[Online; accessed 2020-09-01]
“M-AILABS Speech Multi-lingual Dataset,” https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/ , · 2020
Closest in time.
“Blind C50 estimation from single-channel speech using a convolutional neural network,”
Hannes Gamper, · 2020
Closest in time.
“Data augmentation and loss normalization for deep noise suppression,” 2020
Sebastian Braun and Ivan Tashev, · 2020
Closest in time.