Fetching the paper…
Reading the bibliography…
One of the problems with automated audio captioning (AAC) is the indeterminacy in word selection corresponding to the audio event/scene.
R. Mihalcea and P. Tarau, “TextRank: Bringing Order into Texts,” in
2004
Earlier work this paper cites.
S. Bird, E. Loper and E. Klein, “Natural Language Processing with Python,”
2009
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic Event Detection in Real Life Recordings,” in
2010
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound Technical Demo,” in
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in
2014
Earlier work this paper cites.
D. Barchiesi, D. Giannoulis, D. Stowell, and M. D. Plumbley, “Acoustic Scene Classification: Classifying Environments from the Sounds they Produce,”
2015
Earlier work this paper cites.
M. T. Luong, H. Pham, and C. D. Manning “Effective Approaches to Attention-based Neural Machine Translation,” in
2015
Earlier work this paper cites.
E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic Sound Event Detection using Multi Label Deep Neural Networks in
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for Polyphonic Sound Event Detection,”
2016
Earlier work this paper cites.
K. Drossos, S. Adavanne, and T. Virtanen, “Automated Audio Captioning with Recurrent Neural Networks,” in
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An Ontology and Human-Labeled Dataset for Audio Events,” in
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, DvPlatt, R. A. Saurous, B. Seybold, M. Slaney, R. Weiss, and K. Wilson, “CNN Architectures for LargeScale Audio Classification,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in
2017
Cited alongside, same era.
P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,”
2017
Cited alongside, same era.
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei, “Boosting Image Captioning with Attributes,” in
2017
Cited alongside, same era.
Y. Pan, T. Yao, H. Li, and T. Mei, “Video Captioning with Transferred Semantic Attributes,” in
2017
Cited alongside, same era.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning,” in
2017
Cited alongside, same era.
M. Wu, H. Dinkel, and K. Yu, “Audio Caption: Listen and Tell,” in
2019
Later among the works it cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating Captions for Audios in The Wild,” in
2019
Later among the works it cites.
Y. Koizumi, S. Saito, H. Uematsu, Y. Kawachi, and N. Harada, “Unsupervised Detection of Anomalous Sound based on Deep Learning and the Neyman-Pearson Lemma,”
2019
Later among the works it cites.
W. Boes and H. V. Hamme, “Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Serizel, N. Turpault, H. E. Zadeh, and A. P. Shah, “Large-Scale Weakly Labeled Semi-Supervised Sound Event Detection in Domestic Environments,” in
2018
Cited alongside, same era.
C. Li, W. Xu, S. Li, and S. Gao, “Guiding Generation for Abstractive Text Summarization Based on Key Information Guide Network,” in
2018
Cited alongside, same era.
R. Pasunuru and M. Bansal, “Multi-Reward Reinforced Summarization with Saliency and Entailment,” in
2018
Cited alongside, same era.
S. Gehrmann, Y. Deng, and A. M. Rush, “Bottom-up Abstractive Summarization,” in
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding by Generative Pre-Training,”
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering,” in
2018
Cited alongside, same era.
S. Ikawa and K. Kashino, “Neural Audio Captioning based on Conditional Sequence-to-Sequence Model,” in
2019
Cited alongside, same era.
K. Nishida, I. Saito, K. Nishida, K. Shinoda, A. Otsuka, H. Asano and J. Tomita, “Multi-style Generative Reading Comprehension,” in
2019
Later among the works it cites.
S. Lipping, K. Drossos, and T. Virtanen, “Crowdsourcing a Dataset of Audio Captions,” in
2019
Later among the works it cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An Audio Captioning Dataset,” in
2020
Closest in time.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The NTT DCASE2020 Challenge Task 6 System: Automated Audio Captioning with Keywords and Sentence Length Estimation,” in
2020
Closest in time.
K. Imoto, N. Tonami, Y. Koizumi, M. Yasuda, R. Yamanishi, and Y. Yamashita, “Sound Event Detection By Multitask Learning of Sound Events and Scenes with Soft Scene Labels,” in
2020
Closest in time.
2020
Closest in time.