Fetching the paper…
Reading the bibliography…
Music performances are representative scenarios for audio-visual modeling.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Russ R. Salakhutdinov. 2012 · 2012
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederick P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. 2018 · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018 · 2018
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019 · 2019
Earlier work this paper cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba. 2019 · 2019
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020 · 2020
Earlier work this paper cites.
Temporal reasoning via audio question answering
Haytham M. Fayek and Justin Johnson. 2020 · 2020
Earlier work this paper cites.
Knowit vqa: Answering knowledge-based questions about videos
Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima. 2020 · 2020
Earlier work this paper cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021 · 2021
Cited alongside, same era.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim. 2021 · 2021
Cited alongside, same era.
Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022 · 2022
Cited alongside, same era.
What is missing in deep music generation? a study of repetition and structure in popular music
Shuqi Dai, Huiran Yu, and Roger B Dannenberg. 2022 · 2022
Cited alongside, same era.
Learning to answer questions in dynamic audio-visual scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu. 2022 · 2022
Musecoco: Generating symbolic music from text
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian. 2023 · 2023
Later among the works it cites.
Vlc-bert: visual question answering with contextualized commonsense knowledge
Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao, and Vered Shwartz. 2023 · 2023
Later among the works it cites.
Attention-based methods for audio question answering
Parthasaarathy Sudarsanam and Tuomas Virtanen. 2023 · 2023
Later among the works it cites.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2023 · 2023
Later among the works it cites.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Clotho-aqa: A crowdsourced dataset for audio question answering
Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, and Tuomas Virtanen. 2022 · 2022
Cited alongside, same era.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. 2022 · 2022
Cited alongside, same era.
Avqa: A dataset for audio-visual question answering on videos
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. 2022 · 2022
Cited alongside, same era.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al. 2023 · 2023
Cited alongside, same era.
Av-maskenhancer: Enhancing video representations through audio-visual masked autoencoder
Xingjian Diao, Ming Cheng, and Shitong Cheng. 2023 · 2023
Cited alongside, same era.
Universal source separation with weakly labelled data
Qiuqiang Kong, Ke Chen, Haohe Liu, Xingjian Du, Taylor Berg-Kirkpatrick, Shlomo Dubnov, and Mark D. Plumbley. 2023 · 2023
Cited alongside, same era.
Vision transformers are parameter-efficient audio-visual learners
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius. 2023 · 2023
Cited alongside, same era.
Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks
Haoyi Duan, Yan Xia, Zhou Mingze, Li Tang, Jieming Zhu, and Zhou Zhao. 2024 · 2024
Later among the works it cites.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2024 · 2024
Later among the works it cites.
Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering
Xiulong Liu, Zhikang Dong, and Peng Zhang. 2024 · 2024
Later among the works it cites.
Unsupervised llm adaptation for question answering
Kuniaki Saito, Kihyuk Sohn, Chen-Yu Lee, and Yoshitaka Ushiku. 2024 · 2024
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Later among the works it cites.
Healthq: Unveiling questioning capabilities of llm chains in healthcare conversations
Ziyu Wang, Hao Li, Di Huang, and Amir M. Rahmani. 2024 · 2024
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. 2024 · 2024
Later among the works it cites.