Fetching the paper…
Reading the bibliography…
Data-driven approaches hold promise for audio captioning.
G. C. Tomas Mikolov, Kai Chen, “Efficient estimation of word representations in vector space,” in Proc. Int. Conf. Learn. Represent. , 2013
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in Proc. European Conf. Comput. Vision , 2014, pp. 740–755
2014
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in Proc. IEEE Comput. Soc. Conf. Comput. Vision Pattern Recognit. , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in Proc. European Conf. Comput. Vision , 2016, pp. 382–398
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in Proc. IEEE Workshop Appl. Signal Process. Audio Acoust. , 2017, pp. 374–378
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “AudioSet: An ontology and human-labeled dataset for audio events,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 776–780
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of SPIDEr,” in Proc. IEEE Int. Conf. Comput. Vision , 2017, pp. 873–881
2017
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 41, no. 2, pp. 423–443, 2018
2018
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. N. Am. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol. , 2019, pp. 119–132
2019
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 736–740
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE-ACM Trans. Audio Speech Lang. , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
X. Liu, X. Mei, Q. Huang, J. Sun, J. Zhao, H. Liu, M. D. Plumbley, V. Kilic, and W. Wang, “Leveraging pre-trained BERT for audio captioning,” in Proc. European Signal Proces. Conf. , 2022, pp. 1145–1149
2022
Later among the works it cites.
F. Xiao, J. Guan, H. Lan, Q. Zhu, and W. Wang, “Local information assisted attention-free decoder for audio captioning,” IEEE Signal Process. Lett. , vol. 29, pp. 1604–1608, 2022
2022
Later among the works it cites.
Z. Zhou, Z. Zhang, X. Xu, Z. Xie, M. Wu, and K. Q. Zhu, “Can audio captions be evaluated with image caption metrics?” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 981–985
2022
Later among the works it cites.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
X. Liu, Q. Huang, X. Mei, T. Ko, H. Tang, M. D. Plumbley, and W. Wang, “CL4AC: A contrastive loss for audio captioning,” in Proc. Detection Classification Acoust. Scenes Events Workshop , 2021, pp. 196–200
2021
Cited alongside, same era.
W. Yuan, Q. Han, D. Liu, X. Li, and Z. Yang, “The DCASE 2021 challenge task 6 system: Automated audio captioning with weakly supervised pre-traing and word selection methods,” Detection Classification Acoust. Scenes Events Challenge, Tech. Rep., 2021
2021
Cited alongside, same era.
X. Mei, Q. Huang, X. Liu, G. Chen, J. Wu, Y. Wu, J. Zhao, S. Li, T. Ko, H. L. Tang, X. Shao, M. D. Plumbley, and W. Wang, “An encoder-decoder based audio captioning system with transfer and reinforcement learning,” in Proc. Detection Classification Acoust. Scenes Events Workshop , 2021
2021
Cited alongside, same era.
S.-L. Wu, X. Chang, G. Wichern, J.-w. Jung, F. Germain, J. L. Roux, and S. Watanabe, “BEATs-based audio captioning model with INSTRUCTOR embedding supervision and ChatGPT mix-up,” Detection Classification Acoust. Scenes Events Challenge, Tech. Rep., 2023
2023
Closest in time.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” in Proc. Int. Conf. Machin. Learn. , 2023
2023
Closest in time.
F. Xiao, J. Guan, Q. Zhu, and W. Wang, “Graph attention for automated audio captioning,” IEEE Signal Process. Lett. , vol. 30, pp. 413–417, 2023
2023
Closest in time.
F. Xiao, Q. Zhu, H. Lan, W. Wang, and J. Guan, “Ensemble systems with contrastive language-audio pretraining and attention-based audio features for audio captioning and retrieval,” Detection Classification Acoust. Scenes Events Challenge, Tech. Rep., 2023
2023
Closest in time.