Fetching the paper…
Reading the bibliography…
Audio captioning aims to generate text descriptions from environmental sounds.
“Learning internal representations by error propagation,”
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, · 1985
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, · 2002
Earlier work this paper cites.
“ROUGE: A package for automatic evaluation of summaries,”
C.-Y. Lin, · 2004
Earlier work this paper cites.
“METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,”
S. Banerjee and A. Lavie, · 2005
Earlier work this paper cites.
“Microsoft coco captions: Data collection and evaluation server,”
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick, · 2015
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation,”
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, · 2015
Earlier work this paper cites.
“Spice: Semantic propositional image caption evaluation,”
P. Anderson, B. Fernando, M. Johnson, and S. Gould, · 2016
Earlier work this paper cites.
“Automated audio captioning with recurrent neural networks,”
K. Drossos, S. Adavanne, and T. Virtanen, · 2017
Earlier work this paper cites.
“Cnn architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, · 2017
Earlier work this paper cites.
“Improved image captioning via policy gradient optimization of spider,”
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, · 2017
Earlier work this paper cites.
“AudioCaps: Generating captions for audios in the wild,”
C. D. Kim, B. Kim, H. Lee, and G. Kim, · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., · 2019
Cited alongside, same era.
“Dense relational captioning: Triple-stream networks for relationship-based captioning,”
D.-J. Kim, J. Choi, T.-H. Oh, and I. S. Kweon, · 2019
Cited alongside, same era.
“Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach,”
D.-J. Kim, J. Choi, T.-H. Oh, and I. S. Kweon, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, · 2020
Cited alongside, same era.
“Automated audio captioning by fine-tuning bart with audioset tags,”
F. Gontier, R. Serizel, and C. Cerisara, · 2021
Later among the works it cites.
“Investigating local and global information for automated audio captioning with transfer learning,”
X. Xu, H. Dinkel, M. Wu, Z. Xie, and K. Yu, · 2021
Later among the works it cites.
“Audio captioning transformer,”
X. Mei, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, · 2021
Later among the works it cites.
“Bayesian transformer language models for speech recognition,”
B. Xue, J. Yu, J. Xu, S. Liu, S. Hu, Z. Ye, M. Geng, X. Liu, and H. Meng, · 2021
Later among the works it cites.
“Prefix-tuning: Optimizing continuous prompts for generation,”
X. L. Li and P. Liang, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Audio captioning based on combined audio and semantic embeddings,”
A. Ö. Eren and M. Sert, · 2020
Cited alongside, same era.
“Audio captioning based on transformer and pre-trained cnn,”
K. Chen, Y. Wu, Z. Wang, X. Zhang, F. Nian, S. Li, and X. Shao, · 2020
Cited alongside, same era.
“A transformer-based audio captioning model with keyword estimation,”
Y. Koizumi, R. Masumura, K. Nishida, M. Yasuda, and S. Saito, · 2020
Cited alongside, same era.
“Clotho: An audio captioning dataset,”
K. Drossos, S. Lipping, and T. Virtanen, · 2020
Cited alongside, same era.
Y. Koizumi, Y. Ohishi, D. Niizumi, D. Takeuchi, and M. Yasuda, · 2020
Cited alongside, same era.
“BART: ‘, translation, and comprehension,”
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, · 2020
Cited alongside, same era.
“Mpnet: Masked and permuted pre-training for language understanding,”
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu, · 2020
Cited alongside, same era.
R. Mokady, A. Hertz, and A. H. Bermano, · 2021
Later among the works it cites.
“Lightweight speaker recognition in poincaré spaces,”
J. Lee, K. Sung-Bin, S. Kang, and T.-H. Oh, · 2022
Later among the works it cites.
“Automated audio captioning using transfer learning and reconstruction latent space similarity regularization,”
A. Koh, X. Fuzhao, and C. E. Siong, · 2022
Later among the works it cites.
“High-resolution image synthesis with latent diffusion models,”
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, · 2022
Later among the works it cites.
“Connecting the dots between audio and text without parallel data through visual knowledge transfer,”
Y. Zhao, J. Hessel, Y. Yu, X. Lu, R. Zellers, and Y. Choi, · 2022
Later among the works it cites.
“Dense relational image captioning via multi-task triple-stream networks,”
D.-J. Kim, T.-H. Oh, J. Choi, and I. S. Kweon, · 2022
Later among the works it cites.
“Learning audio-video modalities from image captions,”
A. Nagrani, P. H. Seo, B. Seybold, A. Hauth, S. Manen, C. Sun, and C. Schmid, · 2022
Later among the works it cites.