Fetching the paper…
Reading the bibliography…
Automated audio captioning aims to use natural language to describe the content of audio data.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in
2004
Earlier work this paper cites.
A. Lavie and A. Agarwal, “Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments,” in
2007
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of spider,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in
2017
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in
2020
Cited alongside, same era.
H. Wang, B. Yang, Y. Zou, and D. Chong, “Automated audio captioning with temporal attention,” DCASE2020 Challenge, Tech. Rep., 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
2020
Later among the works it cites.
X. Xu, H. Dinkel, M. Wu, Z. Xie, and K. Yu, “Investigating local and global information for automated audio captioning with transfer learning,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
K. Chen, Y. Wu, Z. Wang, X. Zhang, F. Nian, S. Li, and X. Shao, “Audio captioning based on transformer and pre-trained cnn,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
X. Xu, H. Dinkel, M. Wu, and K. Yu, “A crnn-gru based reinforcement learning approach to audio captioning,” in
2020
Cited alongside, same era.
2021
Closest in time.
H. Xu, Z. Zhengyan, D. Ning, G. Yuxian, L. Xiao, H. Yuqi, Q. Jiezhong, Z. Liang, H. Wentao, H. Minlie, J. Qin, L. Yanyan, L. Yang, L. Zhiyuan, L. Zhiwu, Q. Xipeng, S. Ruihua, T. Jie, W. Ji-Rong, Y. Jinhui, Z. W. Xin, and Z. Jun, “Pre-trained models: Past, present and future,” 2021
2021
Closest in time.
X. Xu, Z. Xie, M. Wu, and K. Yu, “The SJTU system for DCASE2021 challenge task 6: Audio captioning based on encoder pre-training and reinforcement learning,” DCASE2021 Challenge, Tech. Rep., July 2021
2021
Closest in time.
X. Mei, Q. Huang, X. Liu, G. Chen, J. Wu, Y. Wu, J. Zhao, S. Li, T. Ko, H. L. Tang, X. Shao, M. D. Plumbley, and W. Wang, “An encoder-decoder based audio captioning system with transfer and reinforcement learning for DCASE challenge 2021 task 6,” DCASE2021 Challenge, Tech. Rep., July 2021
2021
Closest in time.