Fetching the paper…
Reading the bibliography…
Automated audio captioning aims at generating natural language descriptions for given audio clips, not only detecting and classifying sounds, but also summarizing the relationships between audio events.
P. Kishore, R. Salim, W. Todd, and Z. Wei-Jing, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) , 2002, pp. 311–318
2002
Earlier work this paper cites.
L. Chin-Yew, “Rouge: A package for automatic evaluation of summaries,” in Proceedings of the workshop on text summarization branches out , no. 1, 2004, pp. 25–26
2004
Earlier work this paper cites.
A. Lavie and A. Agarwal, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proceedings of the Second Workshop on Statistical Machine Translation , no. June, 2007, pp. 228–23
2007
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
R. Vedantam, C. L. Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR) , vol. 07-12-June, 2015, pp. 4566–4575
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in Proceedings of European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 382–398
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2017, pp. 374–378
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 830–834
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , 2019, pp. 119–132
2019
Earlier work this paper cites.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The ntt dcase2020 challenge task 6 system: Automated audio captioning with keywords and sentence length estimation,” DCASE2020 Challenge, Tech. Rep., June 2020
2020
Cited alongside, same era.
A. Ö. Eren and M. Sert, “Audio captioning based on combined audio and semantic embeddings,” in 2020 IEEE International Symposium on Multimedia (ISM) . IEEE, 2020, pp. 41–48
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Cakır, K. Drossos, and T. Virtanen, “Multi-task regularization based on infrequent classes for audio captioning,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , Tokyo, Japan, November 2020, pp. 6–10
2020
Cited alongside, same era.
X. Xu, H. Dinkel, M. Wu, and K. Yu, “Audio caption in a car setting with a sentence-level loss,” in Proceedings of the International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5
2021
Later among the works it cites.
X. Liu, Q. Huang, X. Mei, T. Ko, H. Tang, M. D. Plumbley, and W. Wang, “Cl4ac: A contrastive loss for audio captioning,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , Barcelona, Spain, November 2021, pp. 196–200
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Dinkel, M. Wu, and K. Yu, “Towards duration robust weakly supervised sound event detection,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 29, pp. 887–900, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 736–740
2020
Cited alongside, same era.
X. Xu, H. Dinkel, M. Wu, Z. Xie, and K. Yu, “Investigating local and global information for automated audio captioning with transfer learning,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 905–909
2021
Cited alongside, same era.
X. Mei, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, “Audio captioning transformer,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , Barcelona, Spain, November 2021, pp. 211–215
2021
Cited alongside, same era.
Z. Ye, H. Wang, D. Yang, and Y. Zou, “Improving the performance of automated audio captioning via integrating the acoustic and semantic information,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , Barcelona, Spain, 2021, pp. 40–44
2021
Cited alongside, same era.
F. Gontier, R. Serizel, and C. Cerisara, “Automated audio captioning by fine-tuning bart with audioset tags,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , Barcelona, Spain, November 2021, pp. 170–174
2021
Cited alongside, same era.
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 366–370
2021
Later among the works it cites.
N. Turpault, “Analyse des problèmatiques liées à la reconnaissance de sons ambiants en environnement réel,” Theses, Université de Lorraine, May 2021. [Online]. Available: https://hal.inria.fr/tel-03304880
2021
Later among the works it cites.
I. Martin and A. Mesaros, “Diversity and bias in audio captioning datasets,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop (DCASE) , 2021, pp. 90–94
2021
Later among the works it cites.
X. Xu, Z. Xie, M. Wu, and K. Yu, “The SJTU system for DCASE2022 challenge task 6: Audio captioning with audio-text retrieval pre-training,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Later among the works it cites.
Z. Zhou, Z. Zhang, X. Xu, Z. Xie, M. Wu, and K. Q. Zhu, “Can audio captions be evaluated with image caption metrics?” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 981–985
2022
Later among the works it cites.