Fetching the paper…
Reading the bibliography…
Models for audio generation are typically trained on hours of recordings.
Gray, A., Markel, J.: Distance measures for speech processing. IEEE Transactions on Acoustics, Speech, and Signal Processing
1976
Earlier work this paper cites.
Keys, R.: Cubic convolution interpolation for digital image processing. IEEE transactions on acoustics, speech, and signal processing
1981
Earlier work this paper cites.
Tzanetakis, G., Cook, P.: Musical genre classification of audio signals. IEEE Transactions on speech and audio processing
2002
Earlier work this paper cites.
Adler, A., Emiya, V., Jafari, M.G., Elad, M., Gribonval, R., Plumbley, M.D.: Audio inpainting. IEEE Transactions on Audio, Speech, and Language Processing
2011
Earlier work this paper cites.
Bahat, Y., Schechner, Y.Y., Elad, M.: Self-content-based audio inpainting. Signal Processing
2015
Earlier work this paper cites.
Eldar, Y.C.: Sampling theory. Cambridge University Press (2015)
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: Bengio, Y., LeCun, Y. (eds.) 3rd International Conference on Learning Representations, ICLR (2015)
2015
Earlier work this paper cites.
Li, C., Wand, M.: Precomputed real-time texture synthesis with Markovian generative adversarial networks. In: European conference on computer vision. pp. 702–716. Springer (2016)
2016
Earlier work this paper cites.
Lostanlen, V., Cella, C.E.: Deep convolutional networks on the pitch spiral for musical instrument recognition. International Society for Music Information Retrieval (ISMIR) (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Oord, A.v.d., Kalchbrenner, N., Espeholt, L., Kavukcuoglu, K., Vinyals, O., Graves, A.: Conditional image generation with PixelCNN decoders. In: NIPS (2016)
2016
Earlier work this paper cites.
Salimans, T., Kingma, D.P.: Weight Normalization: A simple reparameterization to accelerate training of deep neural networks. In: NIPS (2016)
2016
Earlier work this paper cites.
Yu, F., Koltun, V.: Multi-scale context aggregation by dilated convolutions. International Conference on Learning Representations (ICLR) (2016)
2016
Earlier work this paper cites.
Arjovsky, M., Chintala, S., Bottou, L.: Wasserstein generative adversarial networks. In: International conference on machine learning. pp. 214–223. PMLR (2017)
2017
Earlier work this paper cites.
Defferrard, M., Benzi, K., Vandergheynst, P., Bresson, X.: FMA: A dataset for music analysis. In: 18th International Society for Music Information Retrieval Conference. No. CONF (2017)
2017
Earlier work this paper cites.
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., Courville, A.: Improved training of Wasserstein GANs. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. pp. 5769–5779 (2017)
2017
Earlier work this paper cites.
Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1125–1134 (2017)
2017
Earlier work this paper cites.
Ito, K., Johnson, L.: The LJ speech dataset
2017
Earlier work this paper cites.
Kuleshov, V., Enam, S.Z., Ermon, S.: Audio super-resolution using neural nets. In: ICLR (Workshop Track) (2017)
2017
Earlier work this paper cites.
Manilow, E., Pardo, B.: Leveraging repetition to do audio imputation. In: 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). pp. 309–313. IEEE (2017)
2017
Earlier work this paper cites.
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S., Sotelo, J., Courville, A.C., Bengio, Y.: SampleRNN: An unconditional end-to-end neural audio generation model. In: 5th International Conference on Learning Representations, ICLR (2017)
2017
Earlier work this paper cites.
Pascual, S., Bonafonte, A., Serrà, J.: SEGAN: Speech Enhancement Generative Adversarial Network. Proc. Interspeech 2017 pp. 3642–3646 (2017)
2017
Earlier work this paper cites.
Arik, S., Chen, J., Peng, K., Ping, W., Zhou, Y.: Neural voice cloning with a few samples. In: Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., Garnett, R. (eds.) Advances in Neural Information Processing Systems. vol. 31. Curran Associates, Inc. (2018)
2018
Cited alongside, same era.
Arık, S.Ö., Jun, H., Diamos, G.: Fast spectrogram inversion using multi-head convolutional neural networks. IEEE Signal Processing Letters
2018
Cited alongside, same era.
Donahue, C., McAuley, J., Puckette, M.: Adversarial audio synthesis. In: International Conference on Learning Representations (2018)
2018
Cited alongside, same era.
Engel, J., Agrawal, K.K., Chen, S., Gulrajani, I., Donahue, C., Roberts, A.: GANSynth: Adversarial neural audio synthesis. In: International Conference on Learning Representations (2018)
2018
Cited alongside, same era.
Zhousl16: solo audio
2019
Later among the works it cites.
Bińkowski, M., Donahue, J., Dieleman, S., Clark, A., Elsen, E., Casagrande, N., Cobo, L.C., Simonyan, K.: High Fidelity Speech Synthesis with Adversarial Networks. In: International Conference on Learning Representations (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G., Lockhart, E., Cobo, L., Stimberg, F., et al.: Parallel WaveNet: Fast high-fidelity speech synthesis. In: International conference on machine learning. pp. 3918–3926. PMLR (2018)
2018
Cited alongside, same era.
Perraudin, N., Holighaus, N., Majdak, P., Balazs, P.: Inpainting of long audio segments with similarity graphs. IEEE/ACM Transactions on Audio, Speech, and Language Processing
2018
Cited alongside, same era.
Ping, W., Peng, K., Chen, J.: ClariNet: Parallel wave generation in end-to-end text-to-speech. In: International Conference on Learning Representations (2018)
2018
Cited alongside, same era.
Wang, M., Wu, Z., Kang, S., Wu, X., Jia, J., Su, D., Yu, D., Meng, H.: Speech super-resolution using parallel WaveNet. In: 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP). pp. 260–264. IEEE (2018)
2018
Cited alongside, same era.
Birnbaum, S., Kuleshov, V., Enam, Z., Koh, P.W.W., Ermon, S.: Temporal FiLM: Capturing long-range sequence dependencies with feature-wise modulations. In: Advances in Neural Information Processing Systems (2019)
2019
Cited alongside, same era.
Chandna, P., Blaauw, M., Bonada, J., Gómez, E.: Wgansing: A multi-voice singing voice synthesizer based on the Wasserstein-GAN. In: 2019 27th European Signal Processing Conference (EUSIPCO). pp. 1–5. IEEE (2019)
2019
Cited alongside, same era.
Engel, J., Gu, C., Roberts, A., et al.: DDSP: Differentiable digital signal processing. In: International Conference on Learning Representations (2019)
2019
Cited alongside, same era.
Gupta, A., Shillingford, B., Assael, Y., Walters, T.C.: Speech bandwidth extension with WaveNet. In: 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). pp. 205–208. IEEE (2019)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Hertz, A., Hanocka, R., Giryes, R., Cohen-Or, D.: Deep geometric texture synthesis. ACM Transactions on Graphics (TOG)
2020
Later among the works it cites.
Kim, J., Kim, S., Kong, J., Yoon, S.: Glow-TTS: A generative flow for text-to-speech via monotonic alignment search. Advances in Neural Information Processing Systems
2020
Later among the works it cites.
Lagrange, M., Gontier, F.: Bandwidth extension of musical audio signals with no side information using dilated convolutional neural networks. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 801–805. IEEE (2020)
2020
Later among the works it cites.
Li, Y., Gfeller, B., Tagliasacchi, M., Roblek, D.: Learning to denoise historical music. International Society for Music Information Retrieval (ISMIR) (2020)
2020
Later among the works it cites.
Liu, J.Y., Chen, Y.H., Yeh, Y.C., Yang, Y.H.: Unconditional audio generation with generative adversarial networks and cycle regularization. Proc. Interspeech 2020 pp. 1997–2001 (2020)
2020
Later among the works it cites.
Marafioti, A., Majdak, P., Holighaus, N., Perraudin, N.: GACELA-A generative adversarial context encoder for long audio inpainting of music. IEEE Journal of Selected Topics in Signal Processing (2020)
2020
Later among the works it cites.
Michelashvili, M., Wolf, L.: Hierarchical timbre-painting and articulation generation. International Society for Music Information Retrieval (ISMIR) (2020)
2020
Later among the works it cites.
Nercessian, S.: Zero-shot singing voice conversion. In: Proceedings of the International Society for Music Information Retrieval Conference (2020)
2020
Later among the works it cites.
Ping, W., Peng, K., Zhao, K., Song, Z.: WaveFlow: A compact flow-based model for raw audio. In: International Conference on Machine Learning. pp. 7706–7716. PMLR (2020)
2020
Later among the works it cites.
Sisman, B., Li, H.: Generative adversarial networks for singing voice conversion with and without parallel data. In: Speaker Odyssey. pp. 238–244 (2020)
2020
Later among the works it cites.
Sulun, S., Davies, M.E.: On filter generalization for music bandwidth extension using deep neural networks. IEEE Journal of Selected Topics in Signal Processing (2020)
2020
Later among the works it cites.
Wright, A., Välimäki, V.: Perceptual loss function for neural modeling of audio systems. In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 251–255. IEEE (2020)
2020
Later among the works it cites.
2020
Later among the works it cites.
Zhang, Z., Wang, Y., Gan, C., Wu, J., Tenenbaum, J.B., Torralba, A., Freeman, W.T.: Deep audio priors emerge from harmonic convolutional networks. In: International Conference on Learning Representations (ICLR) (2020)
2020
Later among the works it cites.
Chen, Y.H., Wu, D.Y., Wu, T.H., Lee, H.y.: Again-VC: A one-shot voice conversion using activation guidance and adaptive instance normalization. In: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 5954–5958. IEEE (2021)
2021
Closest in time.
Xu, R., Wang, X., Chen, K., Zhou, B., Loy, C.C.: Positional encoding as spatial inductive bias in gans. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 13569–13578 (2021)
2021
Closest in time.