Fetching the paper…
Reading the bibliography…
Existing efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decisionmaking and heavy reliance on scene-text recognition.
Word spotting and recognition with embedded attributes
Jon Almazán, Albert Gordo, Alicia Fornés, and Ernest Valveny · 2014
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Earlier work this paper cites.
Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
Chuang Gan, Yandong Li, Haoxiang Li, Chen Sun, and Boqing Gong · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marcal Rusinol, CV Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Tvqa+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal · 2019
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Iterative answer prediction with pointer-augmented multimodal transformers for textvqa
Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach · 2020
Earlier work this paper cites.
Reducing language biases in visual question answering with visually-grounded question encoder
Gouthaman Kv and Anurag Mittal · 2020
Earlier work this paper cites.
Visual relation grounding in videos
Junbin Xiao, Xindi Shang, Xun Yang, Sheng Tang, and Tat-Seng Chua · 2020
Earlier work this paper cites.
Vidat—ANU CVML video annotation tool
Jiahao Zhang, Stephen Gould, and Itzik Ben-Shabat · 2020
Earlier work this paper cites.
Multimedia intelligence: When multimedia meets artificial intelligence
Wenwu Zhu, Xin Wang, and Wen Gao · 2020
Earlier work this paper cites.
Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition
Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang · 2021
Earlier work this paper cites.
Ruart: A novel text-centered solution for text-based visual question answering
Zan-Xia Jin, Heran Wu, Chun Yang, Fang Zhou, Jingyan Qin, Lei Xiao, and Xu-Cheng Yin · 2021
Earlier work this paper cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen · 2021
Cited alongside, same era.
A first look: Towards explainable textvqa models via visual and textual explanations
Varun Nagaraj Rao, Xingjian Zhen, Karen Hovsepian, and Mingwei Shen · 2021
Cited alongside, same era.
Dualvgr: A dual-visual graph reasoning unit for video question answering
Jianyu Wang, Bing-Kun Bao, and Changsheng Xu · 2021
Cited alongside, same era.
A bilingual, openworld video text dataset and end-to-end video text spotter with transformer
Weijia Wu, Yuanqiang Cai, Debing Zhang, Sibo Wang, Zhuang Li, Jiahong Li, Yejun Tang, and Hong Zhou · 2021
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Cited alongside, same era.
Symmetrical linguistic feature distillation with clip for scene text recognition
Zixiao Wang, Hongtao Xie, Yuxin Wang, Jianjun Xu, Boqiang Zhang, and Yongdong Zhang · 2023
Later among the works it cites.
Contrastive video question answering via video graph transformer
Junbin Xiao, Pan Zhou, Angela Yao, Yicong Li, Richang Hong, Shuicheng Yan, and Tat-Seng Chua · 2023
Later among the works it cites.
Exploring sparse spatial relation in graph inference for text-based vqa
Sheng Zhou, Dan Guo, Jia Li, Xun Yang, and Meng Wang · 2023
Later among the works it cites.
Locate then generate: Bridging vision and language with bounding box for scene-text vqa
Yongxin Zhu, Zhen Liu, Yukang Liang, Xin Li, Hao Liu, Changcun Bao, and Linli Xu · 2023
Later among the works it cites.
Guess: Gradually enriching synthesis for text-driven human motion generation
Xuehao Gao, Yang Yang, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, and Yang Wu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grounding answers for visual questions asked by visually impaired people
Chongyan Chen, Samreen Anjum, and Danna Gurari · 2022
Cited alongside, same era.
Mukea: Multimodal knowledge extraction and accumulation for knowledge-based visual question answering
Yang Ding, Jing Yu, Bang Liu, Yue Hu, Mingxin Cui, and Qi Wu · 2022
Cited alongside, same era.
Equivariant and invariant grounding for video question answering
Yicong Li, Xiang Wang, Junbin Xiao, and Tat-Seng Chua · 2022
Cited alongside, same era.
Invariant grounding for video question answering
Yicong Li, Xiang Wang, Junbin Xiao, Wei Ji, and Tat-Seng Chua · 2022
Cited alongside, same era.
Disentangled capsule routing for fast part-object relational saliency
Yi Liu, Dingwen Zhang, Nian Liu, Shoukun Xu, and Jungong Han · 2022
Cited alongside, same era.
Scene graph refinement network for visual question answering
Tianwen Qian, Jingjing Chen, Shaoxiang Chen, Bo Wu, and Yu-Gang Jiang · 2022
Cited alongside, same era.
Vqa therapy: Exploring answer differences by visually grounding answers
Chongyan Chen, Samreen Anjum, and Danna Gurari · 2023
Cited alongside, same era.
Benchmarking micro-action recognition: Dataset, method, and application
Dan Guo, Kun Li, Bin Hu, Yan Zhang, and Meng Wang · 2024
Closest in time.
Gomatching: A simple baseline for video text spotting via long and short term matching
Haibin He, Maoyuan Ye, Jing Zhang, Juhua Liu, and Dacheng Tao · 2024
Closest in time.
Deep unsupervised part-whole relational visual saliency
Yi Liu, Xiaohui Dong, Dingwen Zhang, and Shoukun Xu · 2024
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Closest in time.
Locate before answering: Answer guided question localization for video question answering
Tianwen Qian, Ran Cui, Jingjing Chen, Pai Peng, Xiaowei Guo, and Yu-Gang Jiang · 2024
Closest in time.
Emotional video captioning with vision-based emotion interpretation network
Peipei Song, Dan Guo, Xun Yang, Shengeng Tang, and Meng Wang · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Closest in time.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat Seng Chua · 2024
Closest in time.
Robust video question answering via contrastive cross-modality representation learning
Xun Yang, Jianming Zeng, Dan Guo, Shanshan Wang, Jianfeng Dong, and Meng Wang · 2024
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2024
Closest in time.
Learning feature semantic matching for spatio-temporal video grounding
Tong Zhang, Hao Fang, Hao Zhang, Jialin Gao, Xiankai Lu, Xiushan Nie, and Yilong Yin · 2024
Closest in time.
Graph pooling inference network for text-based vqa
Sheng Zhou, Dan Guo, Xun Yang, Jianfeng Dong, and Meng Wang · 2024
Closest in time.
Xuehao Gao, Yang Yang, Shaoyi Du, Guo-Jun Qi, and Junwei Han · 2025
Closest in time.
Egotextvqa: Towards egocentric scene-text aware video question answering
Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao · 2025
Closest in time.