Fetching the paper…
Reading the bibliography…
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demands involved.
G. Tzanetakis, G. Essl, and P. Cook, “Automatic musical genre classification of audio signals,” in The International Society for Music Information Retrieval , 2001
2001
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Proc. Text Summarization Branches Out , 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the 22nd ACM International Conference on Multimedia , 2014, pp. 1041–1044
2014
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “CREMA-D: Crowd-sourced emotional multimodal actors dataset,” IEEE Transactions on Affective Computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in 14th European Conference, Amsterdam, The Netherlands , 2016, pp. 382–398
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2017, pp. 374–378
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of SPIDEr,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 873–881
2017
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE , vol. 13, pp. 1–35, 2018
2018
Earlier work this paper cites.
E. J. Humphrey, S. Durand, and B. McFee, “OpenMIC-2018: An open data-set for multiple instrument recognition,” in International Society for Music Information Retrieval Conference , 2018
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 119–132
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized bert pretraining approach,” arXiv preprint:1907.11692 , 2019
2019
Earlier work this paper cites.
H. Xie and T. Virtanen, “Zero-shot audio classification based on class label embeddings,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 264–267
2019
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 736–740
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7871–7880
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9729–9738
2020
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in Proceedings of the Twelfth Language Resources and Evaluation Conference , 2020, pp. 4218–4222
2020
Cited alongside, same era.
I. Martin Morato and A. Mesaros, “Diversity and bias in audio captioning datasets,” in Proceedings of the 6th Workshop on Detection and Classication of Acoustic Scenes and Events , 2021, pp. 90–94
2021
Cited alongside, same era.
——, “Zero-shot audio classification via semantic embeddings,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1233–1242, 2021
2021
Cited alongside, same era.
A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie, “Audio retrieval with natural language queries: A benchmark study,” IEEE Transactions on Multimedia , vol. 25, pp. 2675–2685, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 3339–3354, 2024
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “Prompting large language models with speech recognition abilities,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 13 351–13 355
2024
Closest in time.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 2871–2883, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Wu, M. Terry, and C. J. Cai, “AI Chains: Transparent and controllable human-ai interaction by chaining large language model prompts,” in Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems , 2022, pp. 1–22
2022
Cited alongside, same era.
X. Mei, X. Liu, M. D. Plumbley, and W. Wang, “Automated audio captioning: An overview of recent progress and new challenges,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2022, no. 1, p. 26, 2022
2022
Cited alongside, same era.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 646–650
2022
Cited alongside, same era.
H. Xie, S. Lipping, and T. Virtanen, “Language-based Audio Retrieval Task in DCASE 2022 Challenge,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , 2022, pp. 216–220
2022
Cited alongside, same era.
X. Mei, X. Liu, J. Sun, M. Plumbley, and W. Wang, “On metric learning for audio-text cross-modal retrieval,” in Interspeech 2022 , 2022, pp. 4142–4146
2022
Cited alongside, same era.
S. Lou, X. Xu, M. Wu, and K. Yu, “Audio-text retrieval in context,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 4793–4797
2022
Cited alongside, same era.
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” arXiv preprint:2303.18223 , 2023
2023
Cited alongside, same era.
2024
Closest in time.
J. Bai, H. Yin, M. Wang, D. Shi, W.-S. Gan, J. Chen, and S. Rahardja, “Audiolog: LLMs-powered long audio logging with hybrid token-semantic contrastive learning,” in IEEE International Conference on Multimedia and Expo (ICME) , 2024, pp. 1–6
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Sun, X. Xu, M. Wu, and W. Xie, “Auto-ACD: A large-scale dataset for audio-language representation learning,” in ACM Multimedia 2024 , 2024
2024
Closest in time.
2024
Closest in time.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in Proceedings of the 12th International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
A. Kwak, C. Morrison, D. Bambauer, and M. Surdeanu, “Classify first, and then extract: Prompt chaining technique for information extraction,” in Proceedings of the Natural Legal Language Processing Workshop 2024 , 2024, pp. 303–317
2024
Closest in time.
2024
Closest in time.
X. Xu, Z. Xie, M. Wu, and K. Yu, “Beyond the Status Quo: A contemporary survey of advances and challenges in audio captioning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 95–112, 2024
2024
Closest in time.
J. Kim, J. Jung, J. Lee, and S. H. Woo, “EnCLAP: Combining neural audio codec and audio-text joint embedding for automated audio captioning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 6735–6739
2024
Closest in time.
E. Labbé, T. Pellegrini, and J. Pinquier, “CoNeTTE: An efficient audio captioning system leveraging multiple datasets with task embedding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 3785–3794, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Dinkel, Y. Wang, Z. Yan, J. Zhang, and Y. Wang, “CED: Consistent ensemble distillation for audio tagging,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 291–295
2024
Closest in time.