Fetching the paper…
Reading the bibliography…
Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them.
Optical-model potential in finite nuclei from reid’s hard core interaction
J-P Jeukenne, A Lejeune, and C Mahaux · 1977
Earlier work this paper cites.
A merging–giveway interaction model of cars in a merging section: a game theoretic analysis
Hideyuki Kita · 1999
Earlier work this paper cites.
Synthesis lectures on artificial intelligence and machine learning
Making · 2009
Earlier work this paper cites.
Weighted banzhaf power and interaction indexes through weighted approximations of games
Jean-Luc Marichal and Pierre Mathonet · 2011
Earlier work this paper cites.
Clustering by fast search and find of density peaks
Alex Rodriguez and Alessandro Laio · 2014
Earlier work this paper cites.
Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems
Anupam Datta, Shayak Sen, and Yair Zick · 2016
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Earlier work this paper cites.
Beyond rnns: Positional self-attention with co-attention for video question answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
A course in game theory
Thomas S Ferguson · 2020
Earlier work this paper cites.
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han · 2020
Earlier work this paper cites.
Modality shifting attention network for multi-modal video question answering
Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, and Chang D Yoo · 2020
Earlier work this paper cites.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Cited alongside, same era.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Random shapley forests: cooperative game-based random forests with consistency
Jianyuan Sun, Hui Yu, Guoqiang Zhong, Junyu Dong, Shu Zhang, and Hongchuan Yu · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Align and prompt: Video-and-language pre-training with entity prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi · 2022
Later among the works it cites.
Toward 3d spatial reasoning for human-like text-based visual question answering
Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen · 2022
Later among the works it cites.
Joint learning of object graph and relation graph for visual question answering
Hao Li, Xu Li, Belhal Karimi, Jie Chen, and Mingming Sun · 2022
Later among the works it cites.
Fine-grained semantically aligned vision-language pre-training
Juncheng Li, Xin He, Longhui Wei, Long Qian, Linchao Zhu, Lingxi Xie, Yueting Zhuang, Qi Tian, and Siliang Tang · 2022
Later among the works it cites.
Dynamic clustering network for unsupervised semantic segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feature augmented memory with global attention network for videoqa
Jiayin Cai, Chun Yuan, Cheng Shi, Lei Li, Yangyang Cheng, and Ying Shan · 2021
Cited alongside, same era.
Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models
Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, and Alexander Hauptmann · 2021
Cited alongside, same era.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Look before you speak: Visually contextualized utterances
Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid · 2021
Cited alongside, same era.
Video question answering: a survey of models and datasets
Guanglu Sun, Lili Liang, Tianlin Li, Bo Yu, Meng Wu, and Bolun Zhang · 2021
Cited alongside, same era.
Kehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu, Yian Zhao, Guoli Song, Li Yuan, and Jie Chen · 2022
Later among the works it cites.
Invariant grounding for video question answering
Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, and Tat-Seng Chua · 2022
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Later among the works it cites.
Multilevel hierarchical network with multiscale sampling for video question answering
Min Peng, Chongyang Wang, Yuan Gao, Yu Shi, and Xiang-Dong Zhou · 2022
Later among the works it cites.
Video question answering with iterative video-text co-tokenization
AJ Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S Ryoo, and Anelia Angelova · 2022
Later among the works it cites.
Locate before answering: Answer guided question localization for video question answering
Tianwen Qian, Ran Cui, Jingjing Chen, Pai Peng, Xiaowei Guo, and Yu-Gang Jiang · 2022
Later among the works it cites.
Learning to answer visual questions from web videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Video question answering: Datasets, algorithms and challenges
Yaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li, Weihong Deng, and Tat-Seng Chua · 2022
Later among the works it cites.
Learning multi-agent intention-aware communication for optimal multi-order execution in finance
Yuchen Fang, Zhenggang Tang, Kan Ren, Weiqing Liu, Li Zhao, Jiang Bian, Dongsheng Li, Weinan Zhang, Yong Yu, and Tieyan Liu · 2023
Closest in time.
Diffusionret: Generative text-video retrieval with diffusion model
Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Xiangyang Ji, Chang Liu, Li Yuan, and Jie Chen · 2023
Closest in time.
Multi-granularity interaction simulation for unsupervised interactive segmentation
Kehan Li, Yian Zhao, Zhennan Wang, Zesen Cheng, Peng Jin, Xiangyang Ji, Li Yuan, Chang Liu, and Jie Chen · 2023
Closest in time.
Fits: Fine-grained two-stage training for knowledge-aware question answering
Qichen Ye, Bowen Cao, Nuo Chen, Weiyuan Xu, and Yuexian Zou · 2023
Closest in time.