Fetching the paper…
Reading the bibliography…
In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Audio-visual speech recognition using deep learning
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G Okuno, and Tetsuya Ogata · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Temporal multimodal learning in audiovisual speech recognition
Di Hu, Xuelong Li, et al · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Exploring multimodal video representation for action recognition
Cheng Wang, Haojin Yang, and Christoph Meinel · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Attention-based bidirectional long short-term memory networks for relation classification
Peng Zhou, Wei Shi, Jun Tian, Zhenyu Qi, Bingchen Li, Hongwei Hao, and Bo Xu · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio-visual speaker diarization based on spatiotemporal bayesian fusion
Israel D Gebru, Sileye Ba, Xiaofei Li, and Radu Horaud · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Learning multimodal attention lstm networks for video captioning
Jun Xu, Ting Yao, Yongdong Zhang, and Tao Mei · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Earlier work this paper cites.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Conversational memory network for emotion recognition in dyadic dialogue videos
Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmermann · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Audio-visual person recognition in multimedia data from the iarpa janus program
Gregory Sell, Kevin Duh, David Snyder, Dave Etter, and Daniel Garcia-Romero · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Audio visual scene-aware dialog
Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K Marks, Chiori Hori, Peter Anderson, et al · 2019
Earlier work this paper cites.
Multi-modal sarcasm detection in twitter with hierarchical fusion model
Yitao Cai, Huiyu Cai, and Xiaojun Wan · 2019
Earlier work this paper cites.
Who said that?: Audio-visual speaker diarisation of real-world meetings
Joon Son Chung, Bong-Jin Lee, and Icksang Han · 2019
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Earlier work this paper cites.
Joint student-teacher learning for audio-visual scene-aware dialog
Chiori Hori, Anoop Cherian, Tim K Marks, and Takaaki Hori · 2019
Earlier work this paper cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Beyond rnns: Positional self-attention with co-attention for video question answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Earlier work this paper cites.
Dual-modality seq2seq network for audio-visual event localization
Yan-Bo Lin, Yu-Jhe Li, and Yu-Chiang Frank Wang · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
A simple baseline for audio-visual scene-aware dialog
Idan Schwartz, Alexander G Schwing, and Tamir Hazan · 2019
Cited alongside, same era.
Noise-tolerant audio-visual online person verification using an attention-based neural network fusion
Suwon Shon, Tae-Hyun Oh, and James Glass · 2019
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Cited alongside, same era.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences
Fengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan, and Guosheng Lin · 2021
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun · 2021
Later among the works it cites.
Adamml: Adaptive multi-modal learning for efficient video recognition
Rameswar Panda, Chun-Fu Richard Chen, Quanfu Fan, Ximeng Sun, Kate Saenko, Aude Oliva, and Rogerio Feris · 2021
Later among the works it cites.
Cross-domain first person audio-visual action recognition through relative norm alignment
Mirco Planamente, Chiara Plizzari, Emanuele Alberti, and Barbara Caputo · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Dave: A deep audio-visual embedding for dynamic saliency prediction
Hamed R Tavakoli, Ali Borji, Esa Rahtu, and Juho Kannala · 2019
Cited alongside, same era.
Audio-visual interpretable and controllable video captioning
Yapeng Tian, Chenxiao Guan, Justin Goodman, Marc Moore, and Chenliang Xu · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Cited alongside, same era.
Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis
Dushyant Singh Chauhan, SR Dhanush, Asif Ekbal, and Pushpak Bhattacharyya · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Audio-visual deep neural network for robust person verification
Yanmin Qian, Zhengyang Chen, and Shuai Wang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
A multi-view approach to audio-visual speaker verification
Leda Sarı, Kritika Singh, Jiatong Zhou, Lorenzo Torresani, Nayan Singhal, and Yatharth Saraf · 2021
Later among the works it cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Later among the works it cites.
From semantic categories to fixations: A novel weakly-supervised visual-auditory saliency detection approach
Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, and Hong Qin · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2021
Later among the works it cites.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Learnable irrelevant modality dropout for multimodal action recognition on modality-specific annotated videos
Saghir Alfasly, Jian Lu, Chen Xu, and Yuru Zou · 2022
Later among the works it cites.
Mm-vit: Multi-modal video transformer for compressed video action recognition
Jiawei Chen and Chiu Man Ho · 2022
Later among the works it cites.
Self-supervised representation learning: Introduction, advances, and challenges
Linus Ericsson, Henry Gouk, Chen Change Loy, and Timothy M Hospedales · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Mix and localize: Localizing sound sources in mixtures
Xixi Hu, Ziyang Chen, and Andrew Owens · 2022
Later among the works it cites.
Expectation-maximization contrastive learning for compact video-and-language representations
Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David Clifton, and Jie Chen · 2022
Later among the works it cites.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2022
Later among the works it cites.
Revealing single frame bias for video-and-language learning
Jie Lei, Tamara L Berg, and Mohit Bansal · 2022
Later among the works it cites.
Learning to answer questions in dynamic audio-visual scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid · 2022
Later among the works it cites.
Audio-visual scene-aware dialog and reasoning using audio-visual transformers with joint student-teacher learning
Ankit Shah, Shijie Geng, Peng Gao, Anoop Cherian, Takaaki Hori, Tim K Marks, Jonathan Le Roux, and Chiori Hori · 2022
Later among the works it cites.
Robust self-supervised audio-visual speech recognition
Bowen Shi, Wei-Ning Hsu, and Abdelrahman Mohamed · 2022
Later among the works it cites.
Multimodal sparse transformer network for audio-visual speech recognition
Qiya Song, Bin Sun, and Shutao Li · 2022
Later among the works it cites.
Human action recognition from various data modalities: A review
Zehua Sun, Qiuhong Ke, Hossein Rahmani, Mohammed Bennamoun, Gang Wang, and Jun Liu · 2022
Later among the works it cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
Learning in audio-visual context: A review, analysis, and new perspective
Yake Wei, Di Hu, Yapeng Tian, and Xuelong Li · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Later among the works it cites.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton · 2022
Later among the works it cites.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Audio-adaptive activity recognition across video domains
Yunhua Zhang, Hazel Doughty, Ling Shao, and Cees GM Snoek · 2022
Later among the works it cites.
An empirical study of end-to-end video-language transformers with masked visual modeling
Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu · 2023
Closest in time.
Sem-pos: Grammatically and semantically correct video captioning
Asmar Nadeem, Adrian Hilton, Robert Dawes, Graham Thomas, and Armin Mustafa · 2023
Closest in time.