Fetching the paper…
Reading the bibliography…
In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Aytar, Y., Vondrick, C., and Torralba, A · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Owens, A., Wu, J., McDermott, J. H., Freeman, W. T., and Torralba, A · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R. and Zisserman, A · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R. J., and Wilson, K · 2017
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Korbar, B., Tran, D., and Torresani, L · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360°video
Morgado, P., Nvasconcelos, N., Langlois, T., and Wang, O · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Senocak, A., Oh, T.-H., Kim, J., Yang, M.-H., and Kweon, I. S · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C · 2018
Earlier work this paper cites.
The sound of pixels
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., and Torralba, A · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Zhou, Y., Wang, Z., Fang, C., Bui, T., and Berg, T. L · 2018
Earlier work this paper cites.
AudioCaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Dual-modality seq2seq network for audio-visual event localization
Lin, Y.-B., Li, Y.-J., and Wang, Y.-C. F · 2019
Cited alongside, same era.
Dual attention matching for audio-visual event localization
Wu, Y., Zhu, L., Yan, Y., and Yang, Y · 2019
Cited alongside, same era.
The sound of motions
Zhao, H., Gan, C., Ma, W.-C., and Torralba, A · 2019
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Cited alongside, same era.
Music gesture for visual sound separation
Gan, C., Huang, D., Zhao, H., Tenenbaum, J. B., and Torralba, A · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Later among the works it cites.
X-clip: End-to-end multi-grained contrastive learning for video-text retrieval
Ma, Y., Xu, G., Sun, X., Yan, M., Zhang, J., and Ji, R · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning representations from audio-visual spatial alignment
Morgado, P., Li, Y., and Nvasconcelos, N · 2020
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Tian, Y., Li, D., and Xu, C · 2020
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Cited alongside, same era.
Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing
Lin, Y.-B., Tseng, H.-Y., Lee, H.-Y., Lin, Y.-Y., and Yang, M.-H · 2021
Cited alongside, same era.
Image super-resolution via iterative refinement
Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Du, Y., Chen, Z., Salamon, J., Russell, B., and Owens, A · 2023
Later among the works it cites.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X., and Zhao, Z · 2023
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M. D · 2023
Later among the works it cites.
Weakly-supervised audio-visual segmentation
Mo, S. and Raj, B · 2023
Later among the works it cites.
Audio-visual class-incremental learning
Pian, W., Mo, S., Guo, Y., and Tian, Y · 2023
Later among the works it cites.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
Ruan, L., Ma, Y., Yang, H., He, H., Liu, B., Fu, J., Yuan, N. J., Jin, Q., and Guo, B · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.