Fetching the paper…
Reading the bibliography…
The goal of audio captioning is to translate input audio into its description using natural language.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic Event Detection in Real Life Recordings,” in Proc. Euro. Signal Process. Conf. (EUSIPCO)
2010
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in Proc. Adv. Neural Inf. Process. Syst. (NIPS),
2014
Earlier work this paper cites.
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu. “Learning Fine-grained Image Similarity with Deep Ranking,” in Proc. IEEE Int. Conf. Comput. Vis. Pattern Recognit. (CVPR)
2014
Earlier work this paper cites.
D. Barchiesi, D. Giannoulis, D. Stowell, and M. D. Plumbley, “Acoustic Scene Classification: Classifying Environments from the Sounds they Produce,” IEEE Signal Process. Mag
2015
Earlier work this paper cites.
M. T. Luong, H. Pham, and C. D. Manning “Effective Approaches to Attention-based Neural Machine Translation,” in Proc. Empir. Methods Nat. Lang. Process. (EMNLP),
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in Proc. Int. Conf. Learn. Representations (ICLR)
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A Unified Embedding for Face Recognition and Clustering’ in Proc. IEEE Int. Conf. on Comput. Vis. Pattern Recognit. (CVPR)
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated Audio Captioning with Recurrent Neural Networks,” in Proc. IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA)
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, DvPlatt, R. A. Saurous, B. Seybold, M. Slaney, R. Weiss, and K. Wilson, “CNN Architectures for LargeScale Audio Classification,” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proc. Adv. Neural Inf. Process. Syst. (NIPS)
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2017
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding with Unsupervised Learning,” Tech. rep., OpenAI,
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding by Generative Pre-Training,” https://blog.openai.com/language-unsupervised
2018
Cited alongside, same era.
S. Ikawa and K. Kashino, “Neural Audio Captioning based on Conditional Sequence-to-Sequence Model,” in Proc. Detect. Classif. Acoust. Scenes Events (DCASE) Workshop
2019
Cited alongside, same era.
L. Luo,Y. Xiong, Y. Liu, and X. Sun, “Adaptive Gradient Methods with Dynamic Bound of Learning Rate,” in Proc. Int. Conf. Learn. Representations (ICLR)
2019
Later among the works it cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An Audio Captioning Dataset,” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2020
Closest in time.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The NTT DCASE2020 Challenge Task 6 System: Automated Audio Captioning with Keywords and Sentence Length Estimation,” in Tech. Rep. Detect. Classif. Acoust, Scenes Events Chall
2020
Closest in time.
Y. Koizumi, R. Masumura, K. Nishida, M. Yasuda, and S. Saito, “A Transformer-based Audio Captioning Model with Keyword Estimation,” in Proc. Interspeech,
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating Captions for Audios in The Wild,” in Proc. N. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Tech. (NAACL-HLT)
2019
Cited alongside, same era.
Y. Koizumi, S. Saito, H. Uematsu, Y. Kawachi, and N. Harada, “Unsupervised Detection of Anomalous Sound based on Deep Learning and the Neyman-Pearson Lemma,” IEEE/ACM Tran. Audio, Speech, and Lang. Process
2019
Cited alongside, same era.
J. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, “Look, Listen and Learn More: Design Choices for Deep Audio Embeddings,” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. N. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Tech. (NAACL-HLT)
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. “Language models are unsupervised multitask learners,” Tech. rep., OpenAI,
2019
Cited alongside, same era.
2020
Closest in time.
K. Imoto, N. Tonami, Y. Koizumi, M. Yasuda, R. Yamanishi, and Y. Yamashita, “Sound Event Detection By Multitask Learning of Sound Events and Scenes with Soft Scene Labels,” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2020
Closest in time.
X. Favory, K. Drossos, T. Virtanen, and X. Serra “COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations,” in Workshop Self-superv. Audio Speech
2020
Closest in time.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating Text Generation with BERT,” in Proc. of Int. Conf. Learn. Representations (ICLR)
2020
Closest in time.
Y. Ohishi, A. Kimura, T. Kawanishi, K. Kashino, D. Harwath and J. Glass “Trilingual Semantic Embeddings of Visually Grounded Speech with Self-attention Mechanisms” in Proc. Int. Conf. Acoust. Speech Signal Process. (ICASSP)
2020
Closest in time.
M. Yasuda, Y. Ohishi, Y. Koizumi, and N. Harada, “rossmodal Sound Retrieval based on Specific Target Co-occurrence Denoted with Weak Labels,” in Proc. Interspeech,
2020
Closest in time.