Fetching the paper…
Reading the bibliography…
Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications.
J. Kontio, L. Laaksonen, and P. Alku, “Neural network-based artificial bandwidth expansion of speech,” Transactions on Audio, Speech, and Language Processing , vol. 15, no. 3, pp. 873–881, 2007
2007
Earlier work this paper cites.
R. M. Bittner, J. Salamon, M. Tierney, M. Mauch, C. Cannam, and J. P. Bello, “MedleyDB: A multitrack dataset for annotation-intensive mir research.” in ISMIR , vol. 14, 2014, pp. 155–160
2014
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJSpeech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “MUSDB18-HQ - an uncompressed version of MUSDB18,” Aug 2019
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in NAACL-HLT , 2019, pp. 119–132
2019
Earlier work this paper cites.
S. Hu, B. Zhang, B. Liang, E. Zhao, and S. Lui, “Phase-aware music super-resolution using generative adversarial networks,” INTERSPEECH , pp. 4074–4078, 2020
2020
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
H. Liu, Q. Kong, Q. Tian, Y. Zhao, D. Wang, C. Huang, and Y. Wang, “VoiceFixer: Toward general speech restoration with neural vocoder,” arXiv preprint:2109.13731 , 2021
2021
Cited alongside, same era.
H. Wang and D. Wang, “Towards robust speech super-resolution,” Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2058–2066, 2021
2021
Cited alongside, same era.
J. Lee and S. Han, “NuWave: A diffusion probabilistic model for neural audio upsampling,” arXiv preprint:2104.02321 , 2021
2021
Cited alongside, same era.
N. C. Rakotonirina, “Self-attention for audio super-resolution,” in International Workshop on Machine Learning for Signal Processing . IEEE, 2021
2021
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” International Conference on Learning Representations , 2022
2022
Later among the works it cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” International Conference on Machine Learning , 2023
2023
Closest in time.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” arXiv preprint:2306.05284 , 2023
2023
Closest in time.
S. Lin, B. Liu, J. Li, and X. Yang, “Common diffusion noise schedules and sample steps are flawed,” arXiv preprint:2305.08891 , 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. You, D. Kim, G. Nam, G. Hwang, and G. Chae, “GAN Vocoder: Multi-resolution discriminator is all you need,” arXiv preprint:2103.05236 , 2021
2021
Cited alongside, same era.
H. Liu, W. Choi, X. Liu, Q. Kong, Q. Tian, and D. Wang, “Neural vocoder is all you need for speech super-resolution,” INTERSPEECH , pp. 4227–4231, 2022
2022
Cited alongside, same era.
S. Han and J. Lee, “NUWave 2: A general neural audio upsampling model for various sampling rates,” arXiv preprint:2206.08545 , 2022
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
2023
Closest in time.
I. Pereira, F. Araújo, F. Korzeniowski, and R. Vogl, “MoisesDB: A dataset for source separation beyond 4-stems,” arXiv preprint:2307.15913 , 2023
2023
Closest in time.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” arXiv preprint:2303.17395 , 2023
2023
Closest in time.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al. , “MusicLM: Generating music from text,” arXiv preprint:2301.11325 , 2023
2023
Closest in time.