Fetching the paper…
Reading the bibliography…
The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years.
J. P. Kincaid, R. P. Fishburne Jr, R. L. Rogers, and B. S. Chissom, “Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel,” Defense Technical Information Center, Fort Belvoir, VA, Tech. Rep., 1975
1975
Earlier work this paper cites.
Y. Koizumi, R. Masumura, K. Nishida, M. Yasuda, and S. Saito, “A Transformer-based audio captioning model with keyword estimation,” in Proc. Interspeech . ISCA, 2020, pp. 1977–1981
1981
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, pp. 533–536, 1986
1986
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out . Association for Computational Linguistics, 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , pp. 65–72, 2005
2005
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proceedings of the 21st ACM International Conference on Multimedia , ser. MM ’13. New York, NY, USA: Association for Computing Machinery, 2013, p. 411–412
2013
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 02 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in 22nd ACM International Conference on Multimedia , Orlando, FL, USA, Nov. 2014, pp. 1041–1044
2014
Earlier work this paper cites.
D. Barchiesi, D. Giannoulis, D. Stowell, and M. D. Plumbley, “Acoustic scene classification: Classifying environments from the sounds they produce,” IEEE Signal Processing Magazine , vol. 32, no. 3, pp. 16–34, 2015
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual question answering,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in European Conference on Computer Vision . Springer, 2016, pp. 382–398
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing , New Orleans, LA, 2017
2017
Earlier work this paper cites.
Y. Xu, Q. Huang, W. Wang, P. Foster, S. Sigtia, P. J. B. Jackson, and M. D. Plumbley, “Unsupervised feature learning based on deep models for environmental audio tagging,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 6, pp. 1230–1241, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
Z.-H. Zhou, “A brief introduction to weakly supervised learning,” National Science Review , vol. 5, no. 1, pp. 44–53, 08 2017
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of SPIDEr,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 873–881
2017
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics . Melbourne, Australia: Association for Computational Linguistics, Jul. 2018, pp. 2556–2565
2018
Earlier work this paper cites.
Q. Kong, C. Yu, Y. Xu, T. Iqbal, W. Wang, and M. D. Plumbley, “Weakly labelled AudioSet tagging with attention neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 11, pp. 1791–1802, 2019
2019
Earlier work this paper cites.
Z. Yu, J. Yu, Y. Cui, D. Tao, and Q. Tian, “Deep modular co-attention networks for visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6281–6290
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019, pp. 119–132
2019
Earlier work this paper cites.
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “NoCaps: Novel object captioning at scale,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 8948–8957
2019
Earlier work this paper cites.
S. Lipping, K. Drossos, and T. Virtanen, “Crowdsourcing a dataset of audio captions,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop , 2019, p. 139
2019
Earlier work this paper cites.
M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2019, pp. 830–834
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
H. M. Fayek and J. Johnson, “Temporal reasoning via audio question answering,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, p. 2283–2294, Aug 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao, “Unified vision-language pre-training for image captioning and VQA,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 13 041–13 049
2020
Cited alongside, same era.
X. Li, X. Yin, C. Li, X. Hu, P. Zhang, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, Y. Choi, and J. Gao, “OSCAR: Object-semantics aligned pre-training for vision-language tasks,” European Conference on Computer Vision , 2020
2020
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2020, pp. 736–740
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020
2020
X. Liu, X. Mei, Q. Huang, J. Sun, J. Zhao, H. Liu, M. D. Plumbley, V. Kilic, and W. Wang, “Leveraging pre-trained BERT for audio captioning,” in 30th European Signal Processing Conference . IEEE, 2022, pp. 1145–1149
2022
Later among the works it cites.
S. Deshmukh, B. Elizalde, and H. Wang, “Audio retrieval with WavText5K and CLAP training,” in Proc. Interspeech . ISCA, 2022
2022
Later among the works it cites.
A. Nagrani, P. H. Seo, B. Seybold, A. Hauth, S. Manen, C. Sun, and C. Schmid, “Learning audio-video modalities from image captions,” in European Conference on Computer Vision . Springer, 2022, pp. 407–426
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022, pp. 646–650
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning . PMLR, 2020, pp. 1597–1607
2020
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Jul. 2020, pp. 7871–7880
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in International Conference on Acoustics, Speech, and Signal Processing , 2020
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
A. Mesaros, T. Heittola, T. Virtanen, and M. D. Plumbley, “Sound event detection: A tutorial,” IEEE Signal Processing Magazine , vol. 38, no. 5, pp. 67–83, 2021
2021
Cited alongside, same era.
F. Gontier, R. Serizel, and C. Cerisara, “Automated audio captioning by fine-tuning BART with AudioSet tags,” in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop , Barcelona, Spain, November 2021, pp. 170–174
2021
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in Neural Information Processing Systems , vol. 34, pp. 9694–9705, 2021
2021
Cited alongside, same era.
Z. Ye, Y. Wang, H. Wang, D. Yang, and Y. Zou, “FeatureCut: An adaptive data augmentation for automated audio captioning,” in 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference , 2022, pp. 313–318
2022
Later among the works it cites.
C. Chen, N. Hou, Y. Hu, H. Zou, X. Qi, and E. S. Chng, “Interactive audio-text representation for automated audio captioning with contrastive learning,” in Proc. Interspeech . ISCA, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello, “Wav2CLIP: Learning robust audio representations from CLIP,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022, pp. 4563–4567
2022
Later among the works it cites.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “AudioCLIP: Extending CLIP to image, text and audio,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022, pp. 976–980
2022
Later among the works it cites.
A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie, “Audio retrieval with natural language queries: A benchmark study,” IEEE Transactions on Multimedia , vol. 25, pp. 2675–2685, 2023
2023
Closest in time.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: Textually guided audio generation,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=CYK7RfcOzQ4
2023
Closest in time.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 23–29 Jul 2023, pp. 21 450–21 474
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1720–1733, 2023
2023
Closest in time.
2023
Closest in time.
Y. Wu*, K. Chen*, T. Zhang*, Y. Hui*, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
2023
Closest in time.
S. Doh and J. Nam, “Lp-musiccaps: Llm-based pseudo music captioning,” in 24th International Society for Music Information Retrieval Conference . International Society for Music Information Retrieval Conference, 2023
2023
Closest in time.
2023
Closest in time.
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria, “Text-to-audio generation using instruction guided latent diffusion model,” in Proceedings of the 31st ACM International Conference on Multimedia . New York, NY, USA: Association for Computing Machinery, 2023, p. 3590–3598
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Y. Xin, D. Yang, and Y. Zou, “Improving text-audio retrieval by text-aware attention pooling and prior matrix revised loss,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023, pp. 1–5
2023
Closest in time.
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, “CLAP learning audio concepts from natural language supervision,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Closest in time.
M. Kim, K. Sung-Bin, and T.-H. Oh, “Prefix tuning for automated audio captioning,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2023
2023
Closest in time.
X. Liu, Q. Huang, X. Mei, H. Liu, Q. Kong, J. Sun, S. Li, T. Ko, Y. Zhang, L. H. Tang et al. , “Visually-aware audio captioning with adaptive audio-visual attention,” in Proc. Interspeech . ISCA, 2023
2023
Closest in time.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 23–29 Jul 2023, pp. 5178–5193
2023
Closest in time.
X. Xu, Z. Zhang, Z. Zhou, P. Zhang, Z. Xie, M. Wu, and K. Q. Zhu, “BLAT: Bootstrapping language-audio pre-training based on audioset tag-guided synthetic data,” in Proceedings of the 31st ACM International Conference on Multimedia . New York, NY, USA: Association for Computing Machinery, 2023, p. 2756–2764
2023
Closest in time.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 2871–2883, 2024
2024
Closest in time.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.