Fetching the paper…
Reading the bibliography…
Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually Indicated Sounds,” in
2016
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold,
2017
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The Sound of Pixels,” in
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,”
2019
Earlier work this paper cites.
A. Nichol and P. Dhariwal, “Improved Denoising Diffusion Probabilistic Models,” in
2019
Earlier work this paper cites.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms,” in
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” in
2020
Earlier work this paper cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A Large-scale Audio-Visual Dataset,” in
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in
2021
Earlier work this paper cites.
V. Iashin and E. Rahtu, “Taming Visually Guided Sound Generation,” in
2021
Earlier work this paper cites.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A Versatile Diffusion Model for Audio Synthesis,” in
2021
Earlier work this paper cites.
T. S. Jonathan Ho, “Classifier-Free Diffusion Guidance,” in
2021
Cited alongside, same era.
A. Rouditchenko, A. Boggust, D. Harwath, B. Chen, D. Joshi, S. Thomas, K. Audhkhasi, H. Kuehne, R. Panda, R. Feris, B. Kingsbury, M. Picheny, A. Torralba, and J. Glass, “AVLnet: Learning Audio-Visual Language Representations from Instructional Videos,” in
2021
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Cited alongside, same era.
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis, “MuLan: A Joint Embedding of Music Audio and Natural Language,” in
2022
Cited alongside, same era.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “AudioCLIP: Extending CLIP to Image, Text and Audio,” in
2022
Cited alongside, same era.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: Textually Guided Audio Generation,” in
2023
Closest in time.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,”
2023
Closest in time.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,” in
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “LAION-5B: An open large-scale dataset for training next generation image-text models,” in
2022
Cited alongside, same era.
2022
Cited alongside, same era.
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello, “Wav2CLIP: Learning Robust Audio Representations From CLIP,” in
2022
Cited alongside, same era.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,” in
2023
Cited alongside, same era.
2023
Closest in time.
S. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A Universal Neural Vocoder with Large-Scale Training,” in
2023
Closest in time.
R. Sheffer and Y. Adi, “I Hear Your True Colors: Image Guided Audio Generation,” in
2023
Closest in time.
H.-W. Dong, N. Takahashi, Y. Mitsufuji, J. McAuley, and T. Berg-Kirkpatrick, “CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos,” in
2023
Closest in time.
H.-W. Dong, G. Sigurdsson, C. Tao, J.-Y. Kao, Y.-H. Lin, A. Narayan-Chen, A. Gupta, T. Chung, J. Huang, N. Peng, and W. Zhao, “CLIPSynth: Learning Text-to-audio Synthesis from Videos Using CLIP and Diffusion Models,” in
2023
Closest in time.
S. Pascual, G. Bhattacharya, C. Yeh, J. Pons, and J. Serrà, “Full-band General Audio Synthesis with Score-based Diffusion,” in
2023
Closest in time.