Fetching the paper…
Reading the bibliography…
Audio question answering (AQA) is a multimodal translation task where a system analyzes an audio signal and a natural language question, to generate a desirable natural language answer.
“Clotho: an audio captioning dataset”
Konstantinos Drossos, Samuel Lipping and Tuomas Virtanen · 1910
Earlier work this paper cites.
“On the Stratification of Multi-label Data”
Konstantinos Sechidis, Grigorios Tsoumakas and Ioannis. Vlahavas · 2011
Earlier work this paper cites.
“Freesound Technical Demo”
Frederic Font, Gerard Roma and Xavier Serra · 2013
Earlier work this paper cites.
“A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input”
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
“VQA: Visual Question Answering”
Aishwarya Agrawal et al · 2015
Earlier work this paper cites.
“Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question”
Haoyuan Gao et al · 2015
Earlier work this paper cites.
“Exploring Nearest Neighbor Approaches for Image Captioning”
Jacob Devlin et al · 2015
Earlier work this paper cites.
“Simple Baseline for Visual Question Answering”
Bolei Zhou et al · 2015
Earlier work this paper cites.
“Visual7W: Grounded Question Answering in Images”
Yuke Zhu, Oliver Groth, Michael. Bernstein and Li Fei-Fei · 2016
Cited alongside, same era.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Cited alongside, same era.
“Yin and Yang: Balancing and Answering Binary Visual Questions”
Peng Zhang et al · 2016
Cited alongside, same era.
“Visual question answering: Datasets, algorithms, and future challenges”
Kushal Kafle and Christopher Kanan · 2017
Cited alongside, same era.
“DeepStory: Video Story QA by Deep Embedded Memory Networks”
Kyung min Kim, Min-Oh Heo, Seongho Choi and Byoung-Tak Zhang · 2017
Cited alongside, same era.
“NewsQA: A Machine Comprehension Dataset”
Adam Trischler et al · 2017
“TVQA: Localized, Compositional Video Question Answering”
Jie Lei, Licheng Yu, Mohit Bansal and Tamara. Berg · 2018
Later among the works it cites.
“CLEAR: A Dataset for Compositional Language and Elementary Acoustic Reasoning”, 2018
Jerome Abdelnour, Giampiero Salvi and Jean Rouat · 2018
Later among the works it cites.
“Advances in Pre-Training Distributed Word Representations”
Tomas Mikolov et al · 2018
Later among the works it cites.
“ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering”
Zhou Yu et al · 2019
Later among the works it cites.
“AudioCaps: Generating Captions for Audios in The Wild”
Chris Kim, Byeongchang Kim, Hyunmin Lee and Gunhee Kim · 2019
Later among the works it cites.
“Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Look, Listen and Learn”
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
“Audio Set: An ontology and human-labeled dataset for audio events”
Jort. Gemmeke et al · 2017
Cited alongside, same era.
Jason Cramer, Ho-Hsiang Wu, Justin Salamon and Juan Bello · 2019
Later among the works it cites.
“Recent Advances in Video Question Answering: A Review of Datasets and Methods”
Devshree Patel, Ratnam Parikh and Yesha Shastri · 2020
Later among the works it cites.
“Temporal Reasoning via Audio Question Answering”
Haytham. Fayek and Justin Johnson · 2020
Later among the works it cites.