Fetching the paper…
Reading the bibliography…
In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos.
Multisensory integration: space, time and superadditivity
Nicholas P Holmes and Charles Spence · 2005
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Visual madlibs: Fill in the blank description generation and question answering
Licheng Yu, Eunbyung Park, Alexander C Berg, and Tamara L Berg · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Attention-based bidirectional long short-term memory networks for relation classification
Peng Zhou, Wei Shi, Jun Tian, Zhenyu Qi, Bingchen Li, Hongwei Hao, and Bo Xu · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Image captioning and visual question answering based on attributes and external knowledge
Qi Wu, Chunhua Shen, Peng Wang, Anthony Dick, and Anton Van Den Hengel · 2017
Earlier work this paper cites.
Learning multimodal attention lstm networks for video captioning
Jun Xu, Ting Yao, Yongdong Zhang, and Tao Mei · 2017
Earlier work this paper cites.
Speech-based visual question answering
Ted Zhang, Dengxin Dai, Tinne Tuytelaars, Marie-Francine Moens, and Luc Van Gool · 2017
Earlier work this paper cites.
Video question answering via hierarchical spatio-temporal attention networks
Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Learning conditioned graph structures for interpretable visual question answering
Will Norcliffe-Brown, Efstathios Vafeias, and Sarah Parisot · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
Audio visual scene-aware dialog
Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K Marks, Chiori Hori, Peter Anderson, et al · 2019
Cited alongside, same era.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Cited alongside, same era.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering
Discriminative sounding objects localization via self-supervised audiovisual matching
Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou · 2020
Later among the works it cites.
Multi-modal dense video captioning
Vladimir Iashin and Esa Rahtu · 2020
Later among the works it cites.
Modality shifting attention network for multi-modal video question answering
Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, and Chang D Yoo · 2020
Later among the works it cites.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Later among the works it cites.
Multiple sound sources localization from coarse to fine
Rui Qian, Heinrich Dinkel Di Hu, Mengyue Wu, Ning Xu, and Weiyao Lin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li · 2019
Cited alongside, same era.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Cited alongside, same era.
Beyond rnns: Positional self-attention with co-attention for video question answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Cited alongside, same era.
A simple baseline for audio-visual scene-aware dialog
Idan Schwartz, Alexander G Schwing, and Tamir Hazan · 2019
Cited alongside, same era.
Audio-visual interpretable and controllable video captioning
Yapeng Tian, Chenxiao Guan, Justin Goodman, Marc Moore, and Chenliang Xu · 2019
Cited alongside, same era.
Dual attention matching for audio-visual event localization
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang · 2019
Cited alongside, same era.
Recursive visual sound separation using minus-plus net
Xudong Xu, Bo Dai, and Dahua Lin · 2019
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Later among the works it cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Later among the works it cites.
Describing unseen videos via multi-modal cooperative dialog agents
Ye Zhu, Yu Wu, Yi Yang, and Yan Yan · 2020
Later among the works it cites.
Multi-level attention fusion network for audio-visual event recognition
Mathilde Brousmiche, Jean Rouat, and Stéphane Dupont · 2021
Later among the works it cites.
Localizing visual sounds the hard way
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Later among the works it cites.
Visualvoice: Audio-visual speech separation with cross-modal consistency
Ruohan Gao and Kristen Grauman · 2021
Later among the works it cites.
Class-aware sounding objects localization via audiovisual correspondence
Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, and Ji-Rong Wen · 2021
Later among the works it cites.
Multimodal{qa}: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant · 2021
Later among the works it cites.
Cyclic co-learning of sounding object visual grounding and sound separation
Yapeng Tian, Di Hu, and Chenliang Xu · 2021
Later among the works it cites.
STAR: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B. Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Yu Wu and Yi Yang · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Self-supervised video object segmentation by motion grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie · 2021
Later among the works it cites.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim · 2021
Later among the works it cites.
Positive sample propagation along the audio-visual event line
Jinxing Zhou, Liang Zheng, Yiran Zhong, Shijie Hao, and Meng Wang · 2021
Later among the works it cites.
Visual sound localization in the wild by cross-modal interference erasing
Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, and Xiaowei Zhou · 2022
Closest in time.
Sepfusion: Finding optimal fusion structures for visual sound separation
Dongzhan Zhou, Xinchi Zhou, Di Hu, Hang Zhou, Lei Bai, Ziwei Liu, and Wanli Ouyang · 2022
Closest in time.