Fetching the paper…
Reading the bibliography…
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, dataaugmentation, multi-task learning, and post-processing, for audiocaptioning.
——, “Pharaoh: a beam search decoder for phrase-based statistical machine translation,” in
2004
Earlier work this paper cites.
P. Koehn,
2009
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real life recordings,” in
2010
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, , and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
D. Barchiesi, D. Giannoulis, D. Stowell, and M. D. Plumbley, “Acoustic scene classification: Classifying environments from the sounds they produce,”
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. Weiss, and K. Wilson, “CNN architectures for largescale audio classification,” in
2017
Earlier work this paper cites.
B. Anderson, P.and Fernando, M. Johnson, and S. Gould, “Guided open vocabulary image captioning with constrained beam search,” in
2017
Cited alongside, same era.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of spider,” in
2017
Cited alongside, same era.
Y. Koizumi, S. Saito, H. Uematsu, Y. Kawachi, and N. Harada, “Unsupervised detection of anomalous sound based on deep learning and the neyman–pearson lemma,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond empirical risk minimization,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
H. Seki, T. Hori, S. Watanabe, N. Moritz, and J. Le Roux, “Vectorized beam search for ctc-attention-based speech recognition.” in
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Lipping, K. Drossos, and T. Virtanen, “Crowdsourcing a dataset of audio captions,” in
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. S. Ayhan and P. Berens, “Test-time data augmentation for estimation of heteroscedastic aleatoric uncertainty in deep neural networks,” in
2018
Cited alongside, same era.
——, “Clotho: An audio captioning dataset,” in
2019
Cited alongside, same era.
M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in
2019
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in
2019
Cited alongside, same era.
S. Ikawa and K. Kashino, “Neural audio captioning based on conditional sequence-to-sequence model,” in
2019
Cited alongside, same era.
DCASE2020 Challenge Task 6: Automated Audio Captioning,
Cited in the paper.
Later among the works it cites.
Y. Koizumi, R. Masumura, K. Nishida, M. Yasuda, , and S. Saito, “A transformer-based audio captioning model with keyword estimation,”
2020
Closest in time.
K. Imoto, N. Tonami, Y. Koizumi, M. Yasuda, R. Yamanishi, and Y. Yamashita, “Sound event detection by multitask learning of sound events and scenes with soft scene labels,” in
2020
Closest in time.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The NTT DCASE2020 challenge task 6 system: Automated audio captioning with keywords and sentence length estimation,” DCASE2020 Challenge, Tech. Rep., 2020
2020
Closest in time.