Fetching the paper…
Reading the bibliography…
Audio Question Answering (AQA) constitutes a pivotal task in which machines analyze both audio signals and natural language questions to produce precise natural language answers.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” in proceedings of the International Conference on Empirical Methods in Natural Language Processing , 2016
2016
Earlier work this paper cites.
Z. Chen, H. Zhang, X. Zhang, and L. Zhao, “Quora question pairs,” 2017
2017
Earlier work this paper cites.
P. Rajpurkar, R. Jia, and P. Liang, “Know what you don’t know: Unanswerable questions for squad,” in Annual Meeting of the Association for Computational Linguistics , 2018
2018
Earlier work this paper cites.
J. Abdelnour, G. Salvi, and J. Rouat, “Clear: A dataset for compositional language and elementary acoustic reasoning,” 2019
2019
Earlier work this paper cites.
A. Roberts, C. Raffel, K. Lee, M. Matena, N. Shazeer, P. J. Liu, S. Narang, W. Li, and Y. Zhou, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Google, Tech. Rep., 2019
2019
Earlier work this paper cites.
H. M. Fayek and J. Johnson, “Temporal reasoning via audio question answering,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2283–2294, 2020
2020
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in proceeding of the IEEE International Conference on Acoustics, Speech and Signal Processing , pp. 736–740, 2020
2020
Cited alongside, same era.
R. Puri, R. Spring, M. Shoeybi, M. Patwary, and B. Catanzaro, “Training question answering models from synthetic data,” in proceedings of the International Conference on Empirical Methods in Natural Language Processing , pp. 5811–5826, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
C. Lyu, L. Shang, Y. Graham, J. Foster, X. Jiang, and Q. Liu, “Improving unsupervised question answering via summarization-informed question generation,” in proceedings of the International Conference on Empirical Methods in Natural Language Processing , pp. 4134–4148, 2021
2021
S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen, “Clotho-AQA: A crowdsourced dataset for audio question answering,” in proceedings of the IEEE International European Signal Processing Conference , pp. 1140–1144, 2022
2022
Later among the works it cites.
G. Li, Y. Wei, Y. Tian, C. Xu, J. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in the proceedings of the International Conference on Computer Vision and Pattern Recognition , pp. 19 086–19 096, 2022
2022
Later among the works it cites.
S. Changpinyo, D. Kukliansy, I. Szpektor, X. Chen, N. Ding, and R. Soricut, “All you may need for VQA are image captions,” in proceedings of the International Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pp. 1947–1963, 2022
2022
Later among the works it cites.
S. R. Behera, B. R. Pailla, A. M. Tripathi, M. B. Rathod, and T. Karavadi, “Towards Multi-Lingual Audio Question Answering,” in proceedings of the IEEE International Speech Communication Association , pp. 356–360, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A. Akula, S. Changpinyo, B. Gong, P. Sharma, S.-C. Zhu, and R. Soricut, “CrossVQA: Scalably generating benchmarks for systematically testing VQA generalization,” in proceedings of the Intrenational Conference on Empirical Methods in Natural Language Processing , pp. 2148–2166, 2021
2021
Cited alongside, same era.
2023
Closest in time.
G. Li, Y. Xu, and D. Hu, “Multi-Scale Attention for Audio Question Answering,” in proceedings of the IEEE International Speech Communication Association , 2023, pp. 3442–3446, 2023
2023
Closest in time.