Fetching the paper…
Reading the bibliography…
The scalability of ambient sound generators is hindered by data scarcity, insufficient caption quality, and limited scalability in model architecture.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Meteor: an automatic metric for mt evaluation with high levels of correlation with human judgments
Lavie, A. and Agarwal, A · 2007
Earlier work this paper cites.
Freesound technical demo
Font, F., Roma, G., and Serra, X · 2013
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R. J., and Wilson, K · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Liu, S., Zhu, Z., Ye, N., Guadarrama, S., and Murphy, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2018
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2019
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Earlier work this paper cites.
Clotho: an audio captioning dataset
Drossos, K., Lipping, S., and Virtanen, T · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Earlier work this paper cites.
Improving image captioning with better use of captions
Shi, Z., Zhou, X., Qiu, X., and Zhu, X · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Automated audio captioning by fine-tuning bart with audioset tags
Gontier, F., Serizel, R., and Cerisara, C · 2021
Earlier work this paper cites.
Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning
Lee, S., Chung, J., Yu, Y., Kim, G., Breuel, T., Chechik, G., and Song, Y · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Earlier work this paper cites.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., and Dubnov, S · 2022
Earlier work this paper cites.
Audio retrieval with wavtext5k and clap training
Deshmukh, S., Elizalde, B., and Wang, H · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models, 2022
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Earlier work this paper cites.
Exploring train and test-time augmentations for audio-language learning
Kim, E., Kim, J., Oh, Y., Kim, K., Park, M., Sim, J., Lee, J., and Lee, K · 2022
Earlier work this paper cites.
Bigvgan: A universal neural vocoder with large-scale training
Lee, S.-g., Ping, W., Ginsburg, B., Catanzaro, B., and Yoon, S · 2022
Earlier work this paper cites.
Learning audio-video modalities from image captions
Nagrani, A., Seo, P. H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C · 2022
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Phenaki: Variable length video generation from open domain textual descriptions
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Cited alongside, same era.
Advancing high-resolution video-language representation with large-scale video transcriptions
Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., and Guo, B · 2022
Cited alongside, same era.
Featurecut: An adaptive data augmentation for automated audio captioning
Ye, Z., Wang, Y., Wang, H., Yang, D., and Zou, Y · 2022
Cited alongside, same era.
Merlot reserve: Neural script knowledge through vision and language and sound
Zellers, R., Lu, J., Lu, X., Yu, Y., Zhao, Y., Salehi, M., Kusupati, A., Hessel, J., Farhadi, A., and Choi, Y · 2022
Cited alongside, same era.
Natural language supervision for general-purpose audio representations
Elizalde, B., Deshmukh, S., and Wang, H · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Closest in time.
Recap: Retrieval-augmented audio captioning
Ghosh, S., Kumar, S., Evuru, C. K. R., Duraiswami, R., and Manocha, D · 2024
Closest in time.
Lafma: A latent flow matching model for text-to-audio generation
Guan, W., Wang, K., Zhou, W., Wang, Y., Deng, F., Wang, H., Li, L., Hong, Q., and Qin, Y · 2024
Closest in time.
Audio generation with multiple conditional diffusion model
Guo, Z., Mao, J., Tao, R., Yan, L., Ouchi, K., Liu, H., and Wang, X · 2024
Closest in time.
Photorealistic video generation with diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Fit: Far-reaching interleaved transformers
Chen, T. and Li, L · 2023
Cited alongside, same era.
SDFusion: Multimodal 3d shape completion, reconstruction, and generation
Cheng, Y.-C., Lee, H.-Y., Tulyakov, S., Schwing, A. G., and Gui, L.-Y · 2023
Cited alongside, same era.
Multilingual audio captioning using machine translated data
Cousin, M., Labbé, E., and Pellegrini, T · 2023
Cited alongside, same era.
Pengi: An audio language model for audio tasks
Deshmukh, S., Elizalde, B., Singh, R., and Wang, H · 2023
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H · 2023
Cited alongside, same era.
Text-to-audio generation using instruction guided latent diffusion model
Ghosal, D., Majumder, N., Mehrish, A., and Poria, S · 2023
Cited alongside, same era.
Gupta, A., Yu, L., Sohn, K., Gu, X., Hahn, M., Li, F.-F., Essa, I., Jiang, L., and Lezama, J · 2024
Closest in time.
Ezaudio: Enhancing text-to-audio generation with efficient diffusion transformer
Hai, J., Xu, Y., Zhang, H., Li, C., Wang, H., Elhilali, M., and Yu, D · 2024
Closest in time.
Discriminator-guided cooperative diffusion for joint audio and video generation
Hayakawa, A., Ishii, M., Shibuya, T., and Mitsufuji, Y · 2024
Closest in time.
Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Kong, Z., Goel, A., Badlani, R., Ping, W., Valle, R., and Catanzaro, B · 2024
Closest in time.
Conette: An efficient audio captioning system leveraging multiple datasets with task embedding
Labb, E., Pellegrini, T., Pinquier, J., et al · 2024
Closest in time.
Audio-free prompt tuning for language-audio models
Li, Y., Wang, X., and Liu, H · 2024
Closest in time.
Wavcraft: Audio editing and generation with large language models
Liang, J., Zhang, H., Liu, H., Cao, Y., Kong, Q., Liu, X., Wang, W., Plumbley, M. D., Phan, H., and Benetos, E · 2024
Closest in time.
Audiosr: Versatile audio super-resolution at scale
Liu, H., Chen, K., Tian, Q., Wang, W., and Plumbley, M. D · 2024
Closest in time.
Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization
Majumder, N., Hung, C.-Y., Ghosal, D., Hsu, W.-N., Mihalcea, R., and Poria, S · 2024
Closest in time.
Tavgbench: Benchmarking text to audible-video generation
Mao, Y., Shen, X., Zhang, J., Qin, Z., Zhou, J., Xiang, M., Zhong, Y., and Dai, Y · 2024
Closest in time.
Foleygen: Visually-guided audio generation
Mei, X., Nagaraja, V., Le Lan, G., Ni, Z., Chang, E., Shi, Y., and Chandra, V · 2024
Closest in time.
Snap video: Scaled spatiotemporal transformers for text-to-video synthesis
Menapace, W., Siarohin, A., Skorokhodov, I., Deyneka, E., Chen, T.-S., Kag, A., Fang, Y., Stoliar, A., Ricci, E., Ren, J., et al · 2024
Closest in time.
Soundlocd: An efficient conditional discrete contrastive latent diffusion model for text-to-sound generation
Niu, X., Zhang, J., Walder, C., and Martin, C. P · 2024
Closest in time.
Soundctm: Uniting score-based and consistency models for text-to-sound generation
Saito, K., Kim, D., Shibuya, T., Lai, C.-H., Zhong, Z., Takida, Y., and Mitsufuji, Y · 2024
Closest in time.
Free sound effects
SoundBible · 2024
Closest in time.
Parameter efficient audio captioning with faithful guidance using audio-text shared latent representation
Sridhar, A. K., Guo, Y., Visser, E., and Mahfuz, R · 2024
Closest in time.
Auto-acd: A large-scale dataset for audio-language representation learning
Sun, L., Xu, X., Wu, M., and Xie, W · 2024
Closest in time.
Codi-2: In-context interleaved and interactive any-to-any generation
Tang, Z., Yang, Z., Khademi, M., Liu, Y., Zhu, C., and Bansal, M · 2024
Closest in time.
Vidmuse: A simple video-to-music generation framework with long-short-term modeling
Tian, Z., Liu, Z., Yuan, R., Pan, J., Liu, Q., Tan, X., Chen, Q., Xue, W., and Guo, Y · 2024
Closest in time.
Beyond deepfake images: Detecting ai-generated videos
Vahdati, D. S., Nguyen, T. D., Azizpour, A., and Stamm, M. C · 2024
Closest in time.
Improving audio captioning models with fine-grained audio features, text embedding supervision, and llm mix-up augmentation
Wu, S.-L., Chang, X., Wichern, G., Jung, J.-w., Germain, F., Le Roux, J., and Watanabe, S · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Xing, Y., He, Y., Tian, Z., Wang, X., and Chen, Q · 2024
Closest in time.
Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation
Xue, J., Deng, Y., Gao, Y., and Li, Y · 2024
Closest in time.
Sound-vecaps: Improving audio generation with visual enhanced captions
Yuan, Y., Jia, D., Zhuang, X., Chen, Y., Liu, Z., Chen, Z., Wang, Y., Wang, Y., Liu, X., Kang, X., et al · 2024
Closest in time.
Zero-shot audio captioning using soft and hard prompts
Zhang, Y., Xu, X., Du, R., Liu, H., Dong, Y., Tan, Z.-H., Wang, W., and Ma, Z · 2024
Closest in time.
Cacophony: An improved contrastive audio-text model
Zhu, G., Darefsky, J., and Duan, Z · 2024
Closest in time.
Cogvlm: Visual expert for pretrained language models
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., XiXuan, S., et al · 2025
Closest in time.