Fetching the paper…
Reading the bibliography…
Audio restoration has become increasingly significant in modern society, not only due to the demand for high-quality auditory experiences enabled by advanced playback devices, but also because the growing capabilities of generative audio models necessitate high-fidelity audio.
K. Brandenburg, “Mp3 and aac explained,” in Audio Engineering Society Conference: 17th International Conference: High-Quality Audio Coding . Audio Engineering Society, 1999
1999
Earlier work this paper cites.
M. Dietz, L. Liljeryd, K. Kjorling, and O. Kunz, “Spectral band replication, a novel approach in audio coding,” in Audio Engineering Society Convention 112 . Audio Engineering Society, 2002
2002
Earlier work this paper cites.
E. Larsen and R. M. Aarts, Audio bandwidth extension: application of psychoacoustics, signal processing and loudspeaker design . John Wiley & Sons, 2005
2005
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE transactions on audio, speech, and language processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “Visqol: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, pp. 1–18, 2015
2015
Earlier work this paper cites.
T. Bäckström, Speech coding: with code-excited linear prediction . Springer, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least squares generative adversarial networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2794–2802
2017
Earlier work this paper cites.
I. Loshchilov, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101 , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “Musdb18-hq-an uncompressed version of musdb18,” doi. org/10.5281/zenodo , vol. 3338373, 2019
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 626–630
2019
Cited alongside, same era.
J. Deng, B. Schuller, F. Eyben, D. Schuller, Z. Zhang, H. Francois, and E. Oh, “Exploiting time-frequency patterns with lstm-rnns for low-bitrate audio restoration,” Neural Computing and Applications , vol. 32, no. 4, pp. 1095–1107, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
Y. Luo and J. Yu, “Music source separation with band-split rnn,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1893–1901, 2023
2023
Later among the works it cites.
K. Li, F. Xie, H. Chen, K. Yuan, and X. Hu, “An audio-visual speech separation model inspired by cortico-thalamo-cortical circuits,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.
X. Li, K. Li, Y. Zheng, C. Yan, X. Ji, and W. Xu, “Safeear: Content privacy-preserving audio deepfake detection,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security , 2024, pp. 3585–3599
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Lattner and J. Nistal, “Stochastic restoration of heavily compressed musical audio using generative adversarial networks,” Electronics , vol. 10, no. 11, p. 1349, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
J. Chen, Y. Shi, W. Liu, W. Rao, S. He, A. Li, Y. Wang, Z. Wu, S. Shang, and C. Zheng, “Gesper: A unified framework for general speech restoration,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–2
2023
Cited alongside, same era.
2023
Cited alongside, same era.
E. Moliner, J. Lehtinen, and V. Välimäki, “Solving audio inverse problems with a diffusion model,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
S. Uhlich, G. Fabbro, M. Hirano, S. Takahashi, G. Wichern et al. , “The sound demixing challenge 2023-cinematic demixing track.”
2023
Cited alongside, same era.
Y.-C. Wu, I. D. Gebru, D. Marković, and A. Richard, “Audiodec: An open-source streaming high-fidelity neural audio codec,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
X. Li, J. Ze, C. Yan, Y. Cheng, X. Ji, and W. Xu, “Enrollment-stage backdoor attacks on speaker recognition systems via adversarial ultrasound,” IEEE Internet of Things Journal , vol. 11, no. 8, pp. 13 108–13 124, 2024
2024
Closest in time.
C. Zeng, X. Miao, X. Wang, E. Cooper, and J. Yamagishi, “Joint speaker encoder and neural back-end model for fully end-to-end automatic speaker verification with multiple enrollment utterances,” Computer Speech & Language , vol. 86, p. 101619, 2024
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” in Advances in Neural Information Processing Systems , 2024
2024
Closest in time.
2024
Closest in time.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Closest in time.
K. Li and Y. Luo, “Subnetwork-to-go: Elastic neural network with dynamic training and customizable inference,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 6775–6779
2024
Closest in time.