Fetching the paper…
Reading the bibliography…
Speech restoration aims at restoring high quality speech in the presence of a diverse set of distortions.
A. Erell and M. Weintraub, “Estimation Using Log-spectral-distance Criterion for Noise-robust Speech Recognition,” in ICASSP , 1990, pp. 853–856
1990
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit,” 2017
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: An Open-source Mandarin Speech Corpus and a Speech Recognition Baseline,” in O-COCOSDA , 2017, pp. 1–5
2017
Earlier work this paper cites.
C. K. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun, P. Rana, S. Srinivasan, and J. Gehrke, “The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results,” in INTERSPEECH , 2020, pp. 2492–2496
2020
Earlier work this paper cites.
S. Braun and I. Tashev, “Data Augmentation and Loss Normalization for Deep Noise Suppression,” in International Conference on Speech and Computer , 2020, pp. 79–86
2020
Earlier work this paper cites.
A. Défossez, G. Synnaeve, and Y. Adi, “Real Time Speech Enhancement in the Waveform Domain,” in INTERSPEECH , 2020, pp. 3291–3295
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. Hao, X. Su, R. Horaud, and X. Li, “FullSubNet: A Full-band and Sub-band Fusion Model for Real-time Single-channel Speech Enhancement,” in ICASSP , 2021, pp. 6633–6637
2021
Earlier work this paper cites.
A. Li, W. Liu, X. Luo, G. Yu, C. Zheng, and X. Li, “A Simultaneous Denoising and Dereverberation Framework with Target Decoupling,” in INTERSPEECH , 2021, pp. 2801–2805
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An End-to-end Neural Audio Codec,” IEEE/ACM TASLP , vol. 30, pp. 495–507, 2021
2021
Earlier work this paper cites.
W. A. Jassim, J. Skoglund, M. Chinen, and A. Hines, “WARP-Q: Quality Prediction for Generative Neural Speech Codecs,” in ICASSP , 2021, pp. 401–405
2021
Earlier work this paper cites.
C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A Non-intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,” in ICASSP , 2021, pp. 6493–6497
2021
Earlier work this paper cites.
J. Serrà, J. Pons, and S. Pascual, “SESQA: Semi-supervised Learning for Speech Quality Assessment,” in ICASSP , 2021, pp. 381–385
2021
Earlier work this paper cites.
S. Pascual, J. Serrà, and J. Pons, “Adversarial Auto-encoding for Packet Loss Concealment,” in WASPAA , 2021, pp. 71–75
2021
Earlier work this paper cites.
N. Kandpal, O. Nieto, and Z. Jin, “Music Enhancement via Image Translation and Vocoding,” in ICASSP , 2022, pp. 3124–3128
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Zhao, B. Ma, K. N. Watcharasupat, and W.-S. Gan, “FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement,” in ICASSP , 2022, pp. 9281–9285
2022
Cited alongside, same era.
H. Liu, X. Liu, Q. Kong, Q. Tian, Y. Zhao, D. Wang, C. Huang, and Y. Wang, “VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration,” in INTERSPEECH , 2022, pp. 4232–4236
2022
Cited alongside, same era.
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman, “MaskGIT: Masked Generative Image Transformer,” in CVPR , 2022, pp. 11 315–11 325
2023
Later among the works it cites.
J. Serrà, D. Scaini, S. Pascual, D. Arteaga, J. Pons, J. Breebaart, and G. Cengarle, “Mono-to-stereo Through Parametric Stereo Generation,” ISMIR , 2023
2023
Later among the works it cites.
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo, “VampNet: Music Generation via Masked Acoustic Token Modeling,” ISMIR , 2023
2023
Later among the works it cites.
R. Sheffer and Y. Adi, “I Hear Your True Colors: Image Guided Audio Generation,” in ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
H. Dubey, V. Gopal, R. Cutler, A. Aazami, S. Matusevych, S. Braun, S. E. Eskimez, M. Thakker, T. Yoshioka, H. Gamper et al. , “ICASSP 2022 Deep Noise Suppression Challenge,” in ICASSP , 2022, pp. 9271–9275
2022
Cited alongside, same era.
M. Liu, S. Lv, Z. Zhang, R. Han, X. Hao, X. Xia, L. Chen, Y. Xiao, and L. Xie, “Two-stage Neural Network for ICASSP 2023 Speech Signal Improvement Challenge,” in ICASSP , 2023, pp. 1–2
2023
Cited alongside, same era.
W. Liu, Y. Shi, J. Chen, W. Rao, S. He, A. Li, Y. Wang, and Z. Wu, “Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction,” in INTERSPEECH , 2023, pp. 4044–4048
2023
Cited alongside, same era.
Y. Koizumi, H. Zen, S. Karita, Y. Ding, K. Yatabe, N. Morioka, Y. Zhang, W. Han, A. Bapna, and M. Bacchiani, “Miipher: A Robust Speech Restoration Model Integrating Self-Supervised Speech and Text Representations,” in WASPPA , 2023, pp. 1–5
2023
Cited alongside, same era.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AUDIOGEN: Textually Guided Audio Generation,” in ICLR , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. P. Murphy, W. T. Freeman, M. Rubinstein et al. , “MUSE: Text-To-Image Generation via Masked Generative Transformers,” in ICML , 2023, pp. 4055–4075
2023
Cited alongside, same era.
2023
Later among the works it cites.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech Enhancement and Dereverberation with Diffusion-based Generative Models,” IEEE/ACM TASLP , 2023
2023
Later among the works it cites.
J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation,” IEEE/ACM TASLP , 2023
2023
Later among the works it cites.
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, “Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit,” in ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
2024
Closest in time.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and Controllable Music Generation,” NeurIPS , vol. 36, 2024
2024
Closest in time.
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An Audio Language Model for Audio Tasks,” NeurIPS , vol. 36, 2024
2024
Closest in time.
Z. Wang, X. Zhu, Z. Zhang, Y. Lv, N. Jiang, G. Zhao, and L. Xie, “SELM: Speech Enhancement Using Discrete Tokens and Language Models,” in ICASSP , 2024
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity Audio Compression with Improved RVQGAN,” NeurIPS , vol. 36, 2024
2024
Closest in time.