Fetching the paper…
Reading the bibliography…
Recent years have seen considerable advances in audio synthesis with deep generative models.
J. Nistal, C. Aouameur, S. Lattner, and G. Richard, “VQCPC-GAN: Variable-length adversarial audio synthesis using vector-quantized contrastive predictive coding,” in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , pp. 116–120, ISSN: 1947-1629
1947
Earlier work this paper cites.
F. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, “CROWDMOS: An approach for crowdsourcing mean opinion score studies,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, pp. 2416–2419. [Online]. Available: http://ieeexplore.ieee.org/document/5946971/
2011
Earlier work this paper cites.
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” in The 9th ISCA Speech Synthesis Workshop, Sunnyvale, CA, USA, 13-15 September 2016 . ISCA, p. 125. [Online]. Available: http://www.isca-speech.org/archive/SSW\_2016/abstracts/ssw9\_DS-4\_van\_den\_Oord.html
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Interspeech 2017 . ISCA, pp. 4006–4010. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2017/wang17n_interspeech.html
2017
Earlier work this paper cites.
E. Richardson and Y. Weiss, “On GANs and GMMs,” in Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31. Curran Associates, Inc. [Online]. Available: https://proceedings.neurips.cc/paper/2018/file/0172d289da48c48de8c5ebf3de9f7ee1-Paper.pdf
2018
Earlier work this paper cites.
M. Binkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying MMD GANs,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net. [Online]. Available: https://openreview.net/forum?id=r1lUOzWCW
2018
Earlier work this paper cites.
C. Donahue, J. J. McAuley, and M. S. Puckette, “Adversarial audio synthesis,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net. [Online]. Available: https://openreview.net/forum?id=ByMVTsR5KQ
2019
Cited alongside, same era.
J. H. Engel, K. K. Agrawal, S. Chen, I. Gulrajani, C. Donahue, and A. Roberts, “GANSynth: Adversarial neural audio synthesis,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net. [Online]. Available: https://openreview.net/forum?id=H1xQVn09FX
2019
Cited alongside, same era.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,” in Interspeech 2019 . ISCA, pp. 2350–2354. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2019/kilgour19_interspeech.html
2019
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net. [Online]. Available: https://openreview.net/forum?id=a-xFK8Ymz5J
2021
Later among the works it cites.
J. Nistal, S. Lattner, and G. Richard, “DarkGAN: Exploiting knowledge distillation for comprehensible audio synthesis with GANs,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021, Online, November 7-12, 2021 , J. H. Lee, A. Lerch, Z. Duan, J. Nam, P. Rao, P. v. Kranenburg, and A. Srinivasamurthy, Eds., pp. 484–492. [Online]. Available: https://archives.ismir.net/ismir2021/paper/000060.pdf
2021
Later among the works it cites.
B. Hayes, C. Saitis, and G. Fazekas, “Neural waveshaping synthesis,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021, Online, November 7-12, 2021 , J. H. Lee, A. Lerch, Z. Duan, J. Nam, P. Rao, P. v. Kranenburg, and A. Srinivasamurthy, Eds., pp. 254–261. [Online]. Available: https://archives.ismir.net/ismir2021/paper/000031.pdf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 3617–3621, ISSN: 2379-190X
2019
Cited alongside, same era.
J. H. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable digital signal processing,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net. [Online]. Available: https://openreview.net/forum?id=B1x1ma4tDr
2020
Cited alongside, same era.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with WaveNet autoencoders,” in Proceedings of the 34th International Conference on Machine Learning . PMLR, pp. 1068–1077, ISSN: 2640-3498. [Online]. Available: https://proceedings.mlr.press/v70/engel17a.html
Cited in the paper.
Adam Neely, “Turning BASS into violin (using AI).” [Online]. Available: https://www.youtube.com/watch?v=cIX4y22NWWc
Cited in the paper.
ANDREW HUANG, “Music with artificial intelligence.” [Online]. Available: https://www.youtube.com/watch?v=AaALLWQmCdI
Cited in the paper.
Cited in the paper.
M. Caetano and N. Osaka, “A formal evaluation framework for sound morphing,” in ICMC
Cited in the paper.
I. Ananthabhotla, S. Ewert, and J. A. Paradiso, “Towards a perceptual loss: Using a neural network codec approximation as a loss for generative audio models,” in Proceedings of the 27th ACM International Conference on Multimedia . ACM, pp. 1518–1525. [Online]. Available: https://dl.acm.org/doi/10.1145/3343031.3351148
Cited in the paper.
H. Leder, B. Belke, A. Oeberst, and D. Augustin, “A model of aesthetic appreciation and aesthetic judgments,” vol. 95, no. 4, pp. 489–508. [Online]. Available: https://bpspsychub.onlinelibrary.wiley.com/doi/abs/10.1348/0007126042369811
Cited in the paper.
2021
Later among the works it cites.
S. Rouard and G. Hadjeres, “CRASH: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021, Online, November 7-12, 2021 , J. H. Lee, A. Lerch, Z. Duan, J. Nam, P. Rao, P. v. Kranenburg, and A. Srinivasamurthy, Eds., pp. 579–585. [Online]. Available: https://archives.ismir.net/ismir2021/paper/000072.pdf
2021
Later among the works it cites.
J. Nistal, S. Lattner, and G. Richard, “Comparing representations for audio synthesis using generative adversarial networks,” in 2020 28th European Signal Processing Conference (EUSIPCO) , pp. 161–165, ISSN: 2076-1465
2076
Closest in time.