Fetching the paper…
Reading the bibliography…
Audio captioning is the task of automatically creating a textual description for the contents of a general audio signal.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in
2004
Earlier work this paper cites.
A. Lavie and A. Agarwal, “Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments,” in
2007
Earlier work this paper cites.
A. Graves,
2012
Earlier work this paper cites.
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” in
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2014
Cited alongside, same era.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in
2015
Cited alongside, same era.
2016
Cited alongside, same era.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in
2016
Cited alongside, same era.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in
2017
Cited alongside, same era.
F. Scheidegger, L. Cavigelli, M. Schaffner, A. C. I. Malossi, C. Bekas, and L. Benini, “Impact of temporal subsampling on accuracy and performance in practical video classification,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Lipping, K. Drossos, and T. Virtanen, “Crowdsourcing a dataset of audio captions,” in
2019
Later among the works it cites.
M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in
2019
Later among the works it cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in
2020
Closest in time.