Fetching the paper…
Reading the bibliography…
Audio captioning is the novel task of general audio content description using free text.
“On the stratification of multi-label data,”
K. Sechidis, G. Tsoumakas, and I. Vlahavas, · 2011
Earlier work this paper cites.
“Freesound technical demo,”
F. Font, G. Roma, and X. Serra, · 2013
Earlier work this paper cites.
“From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, · 2014
Earlier work this paper cites.
“Microsoft COCO: common objects in context,”
T.-Y. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, · 2014
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Deep visual-semantic alignments for generating image descriptions,”
A. Karpathy and L. Fei-Fei, · 2017
Cited alongside, same era.
“Automated audio captioning with recurrent neural networks,”
K. Drossos, S. Adavanne, and T. Virtanen, · 2017
Cited alongside, same era.
“AudioSet: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Later among the works it cites.
“Crowdsourcing a dataset of audio captions,”
S. Lipping, K. Drossos, and T. Virtanen, · 2019
Closest in time.
“Audio caption: Listen and tell,”
M. Wu, H. Dinkel, and K. Yu, · 2019
Closest in time.
“AudioCaps: Generating captions for audios in the wild,”
C. D. Kim, B. Kim, H. Lee, and G. Kim, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…