Fetching the paper…
Reading the bibliography…
Conventionally, audio super-resolution models fixed the initial and the target sampling rates, which necessitate the model to be trained for each pair of sampling rates.
E. O. Brigham and R. Morrow, “The fast fourier transform,” IEEE spectrum , vol. 4, no. 12, pp. 63–70, 1967
1967
Earlier work this paper cites.
K. Li, Z. Huang, Y. Xu, and C.-H. Lee, “Dnn-based speech bandwidth expansion and its application to adding high-frequency missing features for automatic speech recognition of narrowband speech,” in INTERSPEECH , 2015, pp. 2578–2582
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International Conference on Machine Learning , 2015, pp. 2256–2265
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit(version 0.92),” 2016. [Online]. Available: https://datashare.ed.ac.uk/handle/10283/3443
2016
Earlier work this paper cites.
V. Kuleshov, S. Z. Enam, and S. Ermon, “Audio super resolution using neural networks,” in Workshop of International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
T. Y. Lim, R. A. Yeh, Y. Xu, M. N. Do, and M. Hasegawa-Johnson, “Time-frequency networks for audio super-resolution,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 646–650
2018
Earlier work this paper cites.
X. Wang, K. Yu, C. Dong, and C. C. Loy, “Recovering realistic texture in image super-resolution by deep spatial feature transform,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 606–615
2018
Earlier work this paper cites.
X. Li, V. Chebiyyam, K. Kirchhoff, and A. Amazon, “Speech audio super-resolution for speech recognition.” in INTERSPEECH , 2019, pp. 3416–3420
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
S. Birnbaum, V. Kuleshov, Z. Enam, P. W. W. Koh, and S. Ermon, “Temporal film: Capturing long-range sequence dependencies with feature-wise modulations.” in Advances in Neural Information Processing Systems , 2019, pp. 10 287–10 298
2019
Cited alongside, same era.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Advances in Neural Information Processing Systems , 2019, pp. 11 918–11 930
2019
Cited alongside, same era.
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu, “Semantic image synthesis with spatially-adaptive normalization,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 2337–2346
2019
Cited alongside, same era.
2021
Later among the works it cites.
K. Zhang, Y. Ren, C. Xu, and Z. Zhao, “WSRGlow: A Glow-Based Waveform Generative Model for Audio Super-Resolution,” in INTERSPEECH , 2021, pp. 1649–1653
2021
Later among the works it cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” in NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications , 2021
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Hou, C. Xu, V. T. Pham, J. T. Zhou, E. S. Chng, and H. Li, “Speaker and phoneme-aware speech bandwidth extension with residual dual-path network,” in INTERSPEECH , 2020, pp. 4064–4068
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems , 2020, pp. 6840–6851
2020
Cited alongside, same era.
L. Chi, B. Jiang, and Y. Mu, “Fast fourier convolution,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 4479–4488
2020
Cited alongside, same era.
J. Lee and S. Han, “Nu-wave: A diffusion probabilistic model for neural audio upsampling,” in INTERSPEECH , 2021, pp. 1634–1638
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “Wavegrad: Estimating gradients for waveform generation,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
D. P. Kingma, T. Salimans, B. Poole, and J. Ho, “Variational diffusion models,” in Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
V.-A. Nguyen, A. H. T. Nguyen, and A. W. H. Khong, “Tunet: A block-online bandwidth extension model based on transformers and self-supervised pretraining,” in 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 161–165
2022
Closest in time.
R. Suvorov, E. Logacheva, A. Mashikhin, A. Remizova, A. Ashukha, A. Silvestrov, N. Kong, H. Goka, K. Park, and V. Lempitsky, “Resolution-robust large mask inpainting with fourier convolutions,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 2149–2159
2022
Closest in time.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” in International Conference on Learning Representations , 2022
2022
Closest in time.