Fetching the paper…
Reading the bibliography…
Automated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress over the last few years.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Comput. , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
G. A. Miller, “Wordnet: a lexical database for english,” Commun. ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
Earlier work this paper cites.
P. Kishore, R. Salim, W. Todd, and Z. Wei-Jing, “BLEU: a Method for Automatic Evaluation of Machine Translation,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
L. Chin-Yew, “Rouge: A package for automatic evaluation of summaries,” in Proc. Workshop Text Summarization Branches Out , no. 1, 2004, pp. 25–26
2004
Earlier work this paper cites.
D. Wang and G. J. Brown, Computational Auditory Scene Analysis: Principles, Algorithms, and Applications . Hoboken, NJ, USA: Wiley-IEEE press, 2006
2006
Earlier work this paper cites.
A. Lavie and A. Agarwal, “METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments,” in Proc. 2nd Workshop Stat. Mach. Transl. , no. June, 2007, pp. 228–23
2007
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in Proc. Eur. Conf. Comput. Vis. , 2010, pp. 15–29
2010
Earlier work this paper cites.
D. Scherer, A. Müller, and S. Behnke, “Evaluation of pooling operations in convolutional architectures for object recognition,” in Proc. Int. Conf. Artif. Neural Netw. , 2010, pp. 92–101
2010
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: data, models and evaluation metrics,” Journal of Artificial Intelligence Research , vol. 47, no. 1, pp. 853–899, 2013
2013
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proc. ACM Int. Conf. Multimedia , 2013, pp. 411–412
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proc. Eur. Conf. Comput. Vis. , 2014, pp. 740–755
2014
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Proc. Int. Conf. Adv. Neural Inf. Process. Syst. , 2015, pp. 1171–1179
2015
Earlier work this paper cites.
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning, “Generating semantically precise scene graphs from textual descriptions for improved image retrieval,” in Proc. Workshop Vis. Lang. , 2015, pp. 70–80
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: Semantic propositional image caption evaluation,” in Proc. Eur. Conf. Comput. Vis. , 2016, pp. 382–398
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017, pp. 776–780
2017
Earlier work this paper cites.
J. Xu, T. Yao, Y. Zhang, and T. Mei, “Learning multimodal attention lstm networks for video captioning,” in Proc. ACM Int. Conf. Multimedia , 2017, pp. 537–545
2017
Earlier work this paper cites.
K. Drossos, A. Sharath, and V. Tuomas, “Automated audio captioning with recurrent neural networks,” in Proc. IEEE Workshop Appl. Signal Process. Audio Acoust. , 2017, pp. 374–378
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Int. Conf. Adv. Neural Inf. Process. Syst. , vol. 30, pp. 5998–6008, 2017
2017
Earlier work this paper cites.
T. Hori, S. Watanabe, Y. Zhang, and W. Chan, “Advances in joint ctc-attention based end-to-end speech recognition with a deep cnn encoder and rnn-lm,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2017, pp. 949–953
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2017
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 7008–7024
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “A multi-device dataset for urban acoustic scene classification,” in Proc. Detection Classification Acoust. Scenes Events , 2018, pp. 9–13
2018
Earlier work this paper cites.
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra, “Diverse beam search for improved description of complex scenes,” in Proc. AAAI Conf. Artif. Intell. , 2018, pp. 7371–7379
2018
Earlier work this paper cites.
M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2019, pp. 830–834
2019
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol , 2019, pp. 119–132
2019
Earlier work this paper cites.
S. Lipping, K. Drossos, and T. Virtanen, “Crowdsourcing a dataset of audio captions,” in Proc. Detection Classification Acoust. Scenes Events , 2019, pp. 139–143
2019
Earlier work this paper cites.
S. Ikawa and K. Kashino, “Neural audio captioning based on conditional sequence-to-sequence model,” in Proc. Detection Classification Acoust. Scenes Events , 2019, pp. 99–103
2019
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in Proc. Int. Conf. Learn. Representations , 2019, pp. 1–16
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Int. Conf. North Amer. Chapter Assoc. Computat. Linguistics , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
N. Li, Z. Chen, and S. Liu, “Meta learning for image captioning,” in Proc. AAAI Conf. Artif. Intell. , vol. 33, no. 1, 2019, pp. 8626–8633
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , pp. 2613–2617, 2019
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proc. Conf. Empirical Methods Natural Lang. Process. , 2019, pp. 3982–3992
2019
Earlier work this paper cites.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating text generation with bert,” in Proc. Int. Conf. Learn. Representations , 2019, pp. 1–43
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Wisdom, E. Tzinis, H. Erdogan, R. Weiss, K. Wilson, and J. Hershey, “Unsupervised sound separation using mixture invariant training,” in Proc. Int. Conf. Adv. Neural Inf. Process. Syst. , vol. 33, 2020, pp. 3846–3857
2020
Earlier work this paper cites.
Y. Huang, L. Lin, S. Ma, X. Wang, H. Liu, Y. Qian, M. Liu, and K. Ouchi, “Guided multi-branch learning systems for sound event detection with sound separation,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 61–65
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Trans. Audio, Speech Lang. Process. , vol. 28, pp. 2880–2894, 2020
2020
Cited alongside, same era.
S. Perez-Castanos, J. Naranjo-Alcazar, P. Zuccarello, and M. Cobos, “Listen carefully and tell: An audio captioning system based on residual learning and gammatone audio representation,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 150–154
2020
Cited alongside, same era.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The ntt dcase2020 challenge task 6 system: Automated audio captioning with keywords and sentence length estimation,” DCASE2020 Challenge, Tech. Rep., 2020
K. Koutini, J. Schlüter, H. Eghbal-zadeh, and G. Widmer, “Efficient training of audio transformers with patchout,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2022, pp. 2753–2757
2022
Closest in time.
Y. Liang, Y. Long, Y. Li, and J. Liang, “Selective pseudo-labeling and class-wise discriminative fusion for sound event detection,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2022, pp. 1496–1500
2022
Closest in time.
T. Kouzelis, G. Bastas, A. Katsamanis, and A. Potamianos, “Efficient audio captioning transformer with patchout and text guidance,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Closest in time.
A. Koh, X. Fuzhao, and C. E. Siong, “Automated audio captioning using transfer learning and reconstruction latent space similarity regularization,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 7722–7726
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Ö. Eren and M. Sert, “Audio captioning based on combined audio and semantic embeddings,” in Proc. IEEE Int. Symp. Multimedia , 2020
2020
Cited alongside, same era.
X. Xu, H. Dinkel, M. Wu, and K. Yu, “A crnn-gru based reinforcement learning approach to audio captioning,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 225–229
2020
Cited alongside, same era.
D. Takeuchi, Y. Koizumi, Y. Ohishi, N. Harada, and K. Kashino, “Effects of word-frequency based pre- and post- processings for audio captioning,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 190–194
2020
Cited alongside, same era.
K. Chen, Y. Wu, Z. Wang, X. Zhang, F. Nian, S. Li, and X. Shao, “Audio captioning based on transformer and pre-trained cnn,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 21–25
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Cakır, K. Drossos, and T. Virtanen, “Multi-task regularization based on infrequent classes for audio captioning,” in Proc. Detection Classification Acoust. Scenes Events , 2020, pp. 6–10
2020
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 736–740
2020
Cited alongside, same era.
2020
Cited alongside, same era.
K. Chen, J. Wang, F. Deng, and X. Wang, “icnn-transformer: An improved cnn-transformer with channel-spatial attention and keyword prediction for automated audio captioning,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2022, pp. 4167–4171
2022
Closest in time.
C. Chen, N. Hou, Y. Hu, H. Zou, X. Qi, and E. S. Chng, “Interactive auido-text representation for automated audio captioning with contrastive learning,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2022, pp. 2773–2777
2022
Closest in time.
F. Xiao, J. Guan, H. Lan, Q. Zhu, and W. Wang, “Local information assisted attention-free decoder for audio captioning,” IEEE Signal Process. Lett. , vol. 29, pp. 1604–1608, 2022
2022
Closest in time.
X. Liu, X. Mei, Q. Huang, J. Sun, J. Zhao, H. Liu, M. D. Plumbley, V. Kilic, and W. Wang, “Leveraging pre-trained bert for audio captioning,” in Proc. IEEE Eur. Assoc. Signal Process. Conf. , 2022, pp. 1145–1149
2022
Closest in time.
Z. Ye, Y. Zou, F. Cui, and Y. Wang, “Automated audio captioning with multi-task learning,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Closest in time.
Z. Zhou, Z. Zhang, X. Xu, Z. Xie, M. Wu, and K. Q. Zhu, “Can audio captions be evaluated with image caption metrics?” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 981–985
2022
Closest in time.
X. Xu, Z. Xie, M. Wu, and K. Yu, “The SJTU system for DCASE2022 challenge task 6: Audio captioning with audio-text retrieval pre-training,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Closest in time.
X. Xu, M. Wu, and K. Yu, “Diversity-controllable and accurate audio captioning based on neural condition,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 971–975
2022
Closest in time.
X. Mei, X. Liu, J. Sun, M. D. Plumbley, and W. Wang, “Diverse audio captioning via adversarial training,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 8882–8886
2022
Closest in time.
A. Koh, S. Tiwari, and C. E. Siong, “Automated audio captioning with epochal difficult captions for curriculum learning,” in Proc. Asia-Pacific Signal Inf. Process. Assoc. Annu. Summit Conf.s , 2022, pp. 1058–1063
2022
Closest in time.
C. Narisetty, E. Tsunoo, X. Chang, Y. Kashiwagi, M. Hentschel, and S. Watanabe, “Joint speech recognition and audio captioning,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 7892–7896
2022
Closest in time.
Z. Ye, Y. Wang, H. Wang, D. Yang, and Y. Zou, “Featurecut: An adaptive data augmentation for automated audio captioning,” in Proc. Asia-Pacific Signal Inf. Process. Assoc. Annu. Summit Conf. , 2022, pp. 313–318
2022
Closest in time.
OpenAI, “Introducing chatgpt,” https://openai.com/blog/chatgpt , 2022
2022
Closest in time.
E. Labbé, T. Pellegrini, and J. Pinquier, “Is my automatic audio captioning system so bad? spider-max: A metric to consider several caption candidates,” in Proc. Detection Classification Acoust. Scenes Events , 2022, pp. 1–5
2022
Closest in time.
X. Mei, X. Liu, H. Liu, J. Sun, M. D. Plumbley, and W. Wang, “Automated audio captioning with keywords guidance,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Closest in time.
P. Primus and G. Widmer, “Cp-jku’s submission to task 6a of the dcase2022 challenge: a bart encoder-decoder for automatic audio captioning trained via the reinforce algorithm and transfer learning,” DCASE2022 Challenge, Tech. Rep., 2022
2022
Closest in time.
D. Petermann, G. Wichern, A. S. Subramanian, Z.-Q. Wang, and J. Le Roux, “Tackling the cocktail fork problem for separation and transcription of real-world soundtracks,” IEEE/ACM Trans. Audio, Speech Lang. Process. , vol. 31, pp. 2592–2605, 2023
2023
Closest in time.
C. Li, Y. Qian, Z. Chen, D. Wang, T. Yoshioka, S. Liu, Y. Qian, and M. Zeng, “Target sound extraction with variable cross-modality clues,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Closest in time.
S.-L. Wu, X. Chang, G. Wichern, J.-w. Jung, F. Germain, J. L. Roux, and S. Watanabe, “Beats-based audio captioning model with instructor embedding supervision and chatgpt mix-up,” DCASE2023 Challenge, Tech. Rep., 2023
2023
Closest in time.
X. Liu, Q. Huang, X. Mei, H. Liu, Q. Kong, J. Sun, S. Li, T. Ko, Y. Zhang, L. H. Tang, M. D. Plumbley, V. Kılıç, and W. Wang, “Visually-aware audio captioning with adaptive audio-visual attention,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2023, pp. 2838–2842
2023
Closest in time.
F. Xiao, J. Guan, Q. Zhu, and W. Wang, “Graph attention for automated audio captioning,” IEEE Signal Process. Lett. , vol. 30, pp. 413–417, 2023
2023
Closest in time.
M. Kim, K. Sung-Bin, and T.-H. Oh, “Prefix tuning for automated audio captioning,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Closest in time.
T. Schaumlöffel, M. G. Vilas, and G. Roig, “Peacs: Prefix encoding for auditory caption synthesis,” DCASE2023 Challenge, Tech. Rep., 2023
2023
Closest in time.
H. Sun, Z. Yan, Y. Wang, H. Dinkel, J. Zhang, and Y. Wang, “Leveraging multi-task training and image retrieval with clap for audio captioning,” DCASE2023 Challenge, Tech. Rep., 2023
2023
Closest in time.
R. Mahfuz, Y. Guo, and E. Visser, “Improving audio captioning using semantic similarity metrics,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Closest in time.
Y. Zhang, H. Yu, R. Du, Z.-H. Tan, W. Wang, Z. Ma, and Y. Dong, “ACTUAL: Audio captioning with caption feature space regularization,” IEEE/ACM Trans. Audio, Speech Lang. Process. , pp. 1–15, 2023
2023
Closest in time.
D. Emmanouilidou, “Investigations in audio captioning: Addressing vocabulary imbalance and evaluating suitability of language-centric performance metrics,” in Proc. IEEE Eur. Assoc. Signal Process. Conf. , 2023
2023
Closest in time.
Z. Xie, X. Xu, M. Wu, and K. Yu, “Enhance temporal relations in audio captioning with sound event detection,” in Proc. ISCA Annu. Conf. Int. Speech Commun. Assoc. , 2023, pp. 4179–4183
2023
Closest in time.
J.-H. Cho, Y.-A. Park, J. Kim, and J.-H. Chang, “Hyu submission for the dcase 2023 task 6a: automated audio captioning model using al-mixgen and synonyms substitution,” DCASE2023 Challenge, Tech. Rep., 2023
2023
Closest in time.
2023
Closest in time.
S. Bhosale, R. Chakraborty, and S. K. Kopparapu, “A novel metric for evaluating audio caption similarity,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Closest in time.
F. Gontier, R. Serizel, and C. Cerisara, “Spice+: Evaluation of automatic audio captioning systems with pre-trained language models,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. , 2023, pp. 1–5
2023
Closest in time.
E. Labbé, T. Pellegrini, and J. Pinquier, “Irit-ups dcase 2023 audio captioning and retrieval system,” DCASE2023 Challenge, Tech. Rep., 2023
2023
Closest in time.