Fetching the paper…
Reading the bibliography…
Generative adversarial networks have recently demonstrated outstanding performance in neural vocoding outperforming best autoregressive and flow-based models.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen, · 2011
Earlier work this paper cites.
“Least squares generative adversarial networks,”
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley, · 2017
Earlier work this paper cites.
“Noisy speech database for training speech enhancement algorithms and tts models,”
Cassia Valentini-Botinhao et al., · 2017
Earlier work this paper cites.
“Wave-u-net: A multi-scale neural network for end-to-end audio source separation,”
Daniel Stoller, Sebastian Ewert, and Simon Dixon, · 2018
Earlier work this paper cites.
“The unreasonable effectiveness of deep features as a perceptual metric,”
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang, · 2018
Earlier work this paper cites.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling, · 2018
Earlier work this paper cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron Courville, · 2019
Earlier work this paper cites.
“Differentiable consistency constraints for improved deep speech enhancement,”
Scott Wisdom, John R Hershey, Kevin Wilson, Jeremy Thorpe, Michael Chinen, Brian Patton, and Rif A Saurous, · 2019
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al., · 2019
Cited alongside, same era.
“Sdr–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, · 2019
Cited alongside, same era.
“Temporal film: Capturing long-range sequence dependencies with feature-wise modulations,”
Sawyer Birnbaum, Volodymyr Kuleshov, Zayd Enam, Pang Wei Koh, and Stefano Ermon, · 2019
Cited alongside, same era.
“Mosnet: Deep learning based objective assessment for voice conversion,”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 2019
Cited alongside, same era.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Cited alongside, same era.
“A two-stage approach to speech bandwidth extension,”
Ju Lin, Yun Wang, Kaustubh Kalgaonkar, Gil Keren, Didi Zhang, and Christian Fuegen, · 2021
Later among the works it cites.
“Towards robust speech super-resolution,”
Heming Wang and Deliang Wang, · 2021
Later among the works it cites.
“Metricgan+: An improved version of metricgan for speech enhancement,”
Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh, Peter Plantinga, Mirco Ravanelli, Xugang Lu, and Yu Tsao, · 2021
Later among the works it cites.
“Real-time speech frequency bandwidth extension,”
Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov, Victor Ungureanu, and Dominik Roblek, · 2021
Later among the works it cites.
“Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
Chandan KA Reddy, Vishak Gopal, and Ross Cutler, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Music source separation in the waveform domain,”
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach, · 2019
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Seanet: A multi-modal speech enhancement network,”
Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, and Dominik Roblek, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
Eesung Kim and Hyeji Seo, · 2021
Later among the works it cites.
“Voicefixer: A unified framework for high-fidelity speech restoration,”
Haohe Liu, Xubo Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, and DeLiang Wang, · 2022
Closest in time.
“Dnsmos p. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
Chandan KA Reddy, Vishak Gopal, and Ross Cutler, · 2022
Closest in time.
“Dual-branch attention-in-attention transformer for single-channel speech enhancement,”
Guochen Yu, Andong Li, Chengshi Zheng, Yinuo Guo, Yutian Wang, and Hui Wang, · 2022
Closest in time.