Fetching the paper…
Reading the bibliography…
Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly differ from those of other types.
R. Likert, “A technique for the measurement of attitudes.” Archives of Psychology , 1932
1932
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE.” Journal of Machine Learning Research , vol. 9, 2008
2008
Earlier work this paper cites.
R. Bresin, A. de Witt, S. Papetti, M. Civolani, and F. Fontana, “Expressive sonification of footstep sounds,” Proceedings of Interactive Sonification Workshop , 2010
2010
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” International Society for Music Information Retrieval Conference , pp. 591–596, 2011
2011
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” arXiv preprint:1312.6114 , 2013
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” International Conference on Learning Representations , 2014
2014
Earlier work this paper cites.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: A literature survey,” Artificial Intelligence Review , vol. 42, pp. 275–293, 2014
2014
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the ACM International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
D. Herremans and E. Chew, “MorpheuS: Automatic music generation with recurrent pattern constraints and tension profiles,” in Proceedings of IEEE TENCON , 2016, pp. 282–285
2016
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” in ISCA Speech Synthesis Workshop , 2016, pp. 125–125
2016
Earlier work this paper cites.
A. M. Lamb, A. G. ALIAS PARTH GOYAL, Y. Zhang, S. Zhang, A. C. Courville, and Y. Bengio, “Professor forcing: A new algorithm for training recurrent networks,” Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJSpeech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “AudioSet: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2017, pp. 776–780
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “FMA: A dataset for music analysis,” in International Society for Music Information Retrieval Conference , 2017
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2017, pp. 131–135
2017
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019, pp. 119–132
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “Wav2Vec: Unsupervised pre-training for speech recognition,” INTERSPEECH , pp. 3465–3469, 2019
2019
Earlier work this paper cites.
T. Klein and M. Nabi, “Learning to answer by learning to ask: Getting the best of GPT-2 and BERT worlds,” arXiv preprint:1911.02365 , 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
J. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable digital signal processing,” International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” Advances in Neural Information Processing Systems , pp. 8067–8077, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
Y. Qu, P. Liu, W. Song, L. Liu, and M. Cheng, “A text generation and prediction system: Pre-training on new corpora using BERT and GPT-2,” in IEEE International Conference on Electronics Information and Emergency Communication , 2020, pp. 323–326
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2020, pp. 721–725
2020
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
H. Liu, Q. Kong, Q. Tian, Y. Zhao, D. Wang, C. Huang, and Y. Wang, “VoiceFixer: Toward general speech restoration with neural vocoder,” arXiv preprint:2109.13731 , 2021
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
V. Iashin and E. Rahtu, “Taming visually guided sound generation,” in British Machine Vision Conference , 2021
2021
Cited alongside, same era.
Z. Chen, X. Tan, K. Wang, S. Pan, D. Mandic, L. He, and S. Zhao, “Infergrad: Improving diffusion models for vocoder by considering inference in training,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022
2022
Later among the works it cites.
H. Liu, X. Liu, Q. Kong, W. Wang, and M. D. Plumbley, “Learning the spectrogram temporal resolution for audio classification,” arXiv preprint:2210.01719 , 2022
2022
Later among the works it cites.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” arXiv preprint:2210.11416 , 2022
2022
Later among the works it cites.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. Mcgrew, I. Sutskever, and M. Chen, “GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models,” in International Conference on Machine Learning , 2022, pp. 16 784–16 804
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2V-Bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in IEEE Automatic Speech Recognition and Understanding Workshop . IEEE, 2021, pp. 244–250
2021
Cited alongside, same era.
Y. Song, J. Sohl-Dickstein, D. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “Byol for audio: Self-supervised learning for general-purpose audio representation,” in International Joint Conference on Neural Networks , 2021
2021
Cited alongside, same era.
S. Doh, M. Won, K. Choi, and J. Nam, “Toward universal text-to-music retrieval,” arXiv preprint:2211.14558 , 2022
2022
Later among the works it cites.
Y. Cao, S. Li, Y. Liu, Z. Yan, Y. Dai, P. S. Yu, and L. Sun, “A comprehensive survey of AI-generated content: A history of generative AI from GAN to ChatGPT,” arXiv preprint:2303.04226 , 2023
2023
Closest in time.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” International Conference on Machine Learning , 2023
2023
Closest in time.
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, “SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,” arXiv preprint:2305.11000 , 2023
2023
Closest in time.
X. Liu, Z. Zhu, H. Liu, Y. Yuan, M. Cui, Q. Huang, J. Liang, Y. Cao, Q. Kong, M. D. Plumbley et al. , “WavJourney: Compositional audio creation with large language models,” arXiv preprint:2307.14335 , 2023
2023
Closest in time.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al. , “MusicLM: Generating music from text,” arXiv preprint:2301.11325 , 2023
2023
Closest in time.
M. W. Lam, Q. Tian, T. Li, Z. Yin, S. Feng, M. Tu, Y. Ji, R. Xia, M. Ma, X. Song et al. , “Efficient neural music generation,” arXiv preprint:2305.15719 , 2023
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1720–1733, 2023
2023
Closest in time.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-to-audio generation with prompt-enhanced diffusion models,” International Conference on Machine Learning , 2023
2023
Closest in time.
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “ImageBind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 180–15 190
2023
Closest in time.
Q. Kong, K. Chen, H. Liu, X. Du, T. Berg-Kirkpatrick, S. Dubnov, and M. D. Plumbley, “Universal source separation with weakly labelled data,” arXiv preprint:2305.07447 , 2023
2023
Closest in time.
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” arXiv preprint:2303.18223 , 2023
2023
Closest in time.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “AudioLM: A language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 42, pp. 2523–2544, 2023
2023
Closest in time.
R. Sheffer and Y. Adi, “I hear your true colors: Image guided audio generation,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank et al. , “Noise2Music: Text-conditioned music generation with diffusion models,” arXiv preprint:2302.03917 , 2023
2023
Closest in time.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” arXiv preprint:2306.05284 , 2023
2023
Closest in time.
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria, “Text-to-audio generation using instruction-tuned LLM and latent diffusion model,” arXiv preprint:2304.13731 , 2023
2023
Closest in time.
X. Liu, H. Liu, Q. Kong, X. Mei, M. D. Plumbley, and W. Wang, “Simple pooling front-ends for efficient audio classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
X. Tan, T. Qin, J. Bian, T.-Y. Liu, and Y. Bengio, “Regeneration learning: A learning paradigm for data generation,” arXiv preprint:2301.08846 , 2023
2023
Closest in time.
Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Lin, A. Ragni, E. Benetos, N. Gyenge et al. , “MERT: Acoustic music understanding model with large-scale self-supervised training,” arXiv preprint:2306.00107 , 2023
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
H.-H. Wu, O. Nieto, J. P. Bello, and J. Salomon, “Audio-text models do not yet leverage natural language,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
X. Tan, Neural Text-to-Speech Synthesis , ser. Artificial Intelligence: Foundations, Theory, and Algorithms. Springer Singapore, 2023
2023
Closest in time.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” arXiv preprint:2303.17395 , 2023
2023
Closest in time.
Z. Chen, N. Kanda, J. Wu, Y. Wu, X. Wang, T. Yoshioka, J. Li, S. Sivasankaran, and S. E. Eskimez, “Speech separation with large-scale self-supervised learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
J. Huang, Y. Ren, R. Huang, D. Yang, Z. Ye, C. Zhang, J. Liu, X. Yin, Z. Ma, and Z. Zhao, “Make-An-Audio 2: Temporal-enhanced text-to-audio generation,” arXiv preprint:2305.18474 , 2023
2023
Closest in time.
F. Schneider, Z. Jin, and B. Schölkopf, “Mousai: Text-to-music generation with long-context latent diffusion,” arXiv preprint:2301.11757 , 2023
2023
Closest in time.