Fetching the paper…
Reading the bibliography…
Automated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures.
B. T. Lowerre, “The harpy speech recognition system.” Ph.D. dissertation, Carnegie Mellon University, USA, 1976
1976
Earlier work this paper cites.
S. Bird, E. Klein, and E. Loper, Natural language processing with Python: analyzing text with the natural language toolkit . ” O’Reilly Media, Inc.”, 2009
2009
Earlier work this paper cites.
2013
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in 2015 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) . Los Alamitos, CA, USA: IEEE Computer Society, jun 2015, pp. 3156–3164
2015
Earlier work this paper cites.
R. Vedantam, C. L. Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in 2015 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” 2016
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in 2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in Computer Vision – ECCV 2016 , B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds. Cham: Springer International Publishing, 2016, pp. 382–398
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated audio captioning with recurrent neural networks,” in 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2017, pp. 374–378
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP) . New Orleans, LA: IEEE, Mar. 2017, pp. 776–780
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) , July 2017
2017
Earlier work this paper cites.
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy, “Improved image captioning via policy gradient optimization of spider,” in IEEE International Conf. on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . IEEE Computer Society, 2017, pp. 873–881
2017
Earlier work this paper cites.
H. Zhang, M. Cissé, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in 6th International Conf. on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conf. Track Proc. OpenReview.net, 2018
2018
Earlier work this paper cites.
P. Izmailov, D. Podoprikhin, T. Garipov, D. P. Vetrov, and A. G. Wilson, “Averaging weights leads to wider optima and better generalization,” in Conf. on Uncertainty in Artificial Intelligence , 2018
2018
Earlier work this paper cites.
A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” 2018
2018
Earlier work this paper cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition , 2018, pp. 4510–4520
2018
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “A multi-device dataset for urban acoustic scene classification,” in Proc. of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018) , November 2018, pp. 9–13
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 119–132
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Interspeech 2019 . ISCA, sep 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in 7th International Conf. on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019
2019
Cited alongside, same era.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in 2020 IEEE International Conf. on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020 . IEEE, 2020, pp. 736–740
2020
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proc. of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 7871–7880
2020
Cited alongside, same era.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, and M. Plumbley, “PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 01 2020
E. Kim, J. Kim, Y. Oh, K. Kim, M. Park, J. Sim, J. Lee, and K. Lee, “Improving audio-language learning with mixgen and multi-level test-time augmentation,” 2022
2022
Later among the works it cites.
I. Shin, Y.-H. Tsai, B. Zhuang, S. Schulter, B. Liu, S. Garg, I. S. Kweon, and K.-J. Yoon, “MM-TTA: Multi-modal test-time adaptation for 3d semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection,” in ICASSP 2022 - 2022 IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 646–650
2022
Later among the works it cites.
T. Kouzelis, G. Bastas, A. Katsamanis, and A. Potamianos, “Efficient audio captioning transformer with patchout and text guidance,” DCASE2022 Challenge, Tech. Rep., July 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” 2020
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, jan 2020
2020
Cited alongside, same era.
D. Takeuchi, Y. Koizumi, Y. Ohishi, N. Harada, and K. Kashino, “Effects of word-frequency based pre- and post- processings for audio captioning,” 2020
2020
Cited alongside, same era.
M. Honnibal, I. Montani, S. Van Landeghem, and A. Boyd, “spaCy: Industrial-strength Natural Language Processing in Python,” 2020
2020
Cited alongside, same era.
I. Martin and A. Mesaros, “Diversity and bias in audio captioning datasets,” in Proc. of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021) , Barcelona, Spain, November 2021, pp. 90–94
2021
Cited alongside, same era.
X. Mei, Q. Huang, X. Liu, G. Chen, J. Wu, Y. Wu, J. Zhao, S. Li, T. Ko, H. L. Tang, X. Shao, M. D. Plumbley, and W. Wang, “An encoder-decoder based audio captioning system with transfer and reinforcement learning,” 2021
2021
Cited alongside, same era.
X. Mei, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, “Audio captioning transformer,” in Proc. of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021) , Barcelona, Spain, November 2021, pp. 211–215
2021
Cited alongside, same era.
F. Gontier, R. Serizel, and C. Cerisara, “Automated audio captioning by fine-tuning bart with audioset tags,” in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021) , Barcelona, Spain, November 2021, pp. 170–174
2021
Cited alongside, same era.
K. Koutini, J. Schlüter, H. Eghbal-zadeh, and G. Widmer, “Efficient Training of Audio Transformers with Patchout,” in Proc. Interspeech 2022 , 2022, pp. 2753–2757
2022
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” 2022
2022
Later among the works it cites.
X. Mei, X. Liu, J. Sun, M. D. Plumbley, and W. Wang, “Diverse audio captioning via adversarial training,” 2022
2022
Later among the works it cites.
Z. Zhou, Z. Zhang, X. Xu, Z. Xie, M. Wu, and K. Q. Zhu, “Can audio captions be evaluated with image caption metrics?” in ICASSP 2022 - 2022 IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 981–985
2022
Later among the works it cites.
I. Martín-Morató, M. Harju, and A. Mesaros, “A summarization approach to evaluating audio captioning,” in Proc. of the 7th Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022) , Nancy, France, November 2022
2022
Later among the works it cites.
E. Labbé, T. Pellegrini, and J. Pinquier, “Is my Automatic Audio Captioning System so Bad? SPIDEr-max: A Metric to Consider Several Caption Candidates,” in Proc. of the 7th Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022) , Nancy, France, November 2022
2022
Later among the works it cites.
2023
Closest in time.
S.-L. Wu, X. Chang, G. Wichern, J.-w. Jung, F. Germain, J. L. Roux, and S. Watanabe, “BEATs-based audio captioning model with INSTRUCTOR embedding supervision and ChatGPT mix-up,” DCASE2023 Challenge, Tech. Rep., May 2023
2023
Closest in time.
H. Su, W. Shi, J. Kasai, Y. Wang, Y. Hu, M. Ostendorf, W. tau Yih, N. A. Smith, L. Zettlemoyer, and T. Yu, “One embedder, any task: Instruction-finetuned text embeddings,” 2023
2023
Closest in time.
T. Pellegrini, I. Khalfaoui-Hassani, E. Labbé, and T. Masquelier, “Adapting a ConvNeXt model to audio classification on AudioSet,” in Accepted to Interspeech . ISCA, sep 2023
2023
Closest in time.
E. Labbé, J. Pinquier, and T. Pellegrini, “Multitask learning in audio captioning: a sentence embedding regression loss acts as a regularizer,” 2023
2023
Closest in time.
E. Labbé, T. Pellegrini, and J. Pinquier, “Irit-ups dcase 2023 audio captioning and retrieval system,” DCASE2023 Challenge, Tech. Rep., May 2023
2023
Closest in time.
M. Kadlčík, A. Hájek, J. Kieslich, and R. Winiecki, “A whisper transformer for audio captioning trained with synthetic captions and transfer learning,” DCASE2023 Challenge, Tech. Rep., May 2023
2023
Closest in time.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML . PMLR, 2023, pp. 28 492–28 518
2023
Closest in time.
F. Gontier, R. Serizel, and C. Cerisara, “SPICE+: Evaluation of Automatic Audio Captioning Systems with Pre-Trained Language Models,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.