Fetching the paper…
Reading the bibliography…
Recent advances in video question answering (VideoQA) offer promising applications, especially in traffic monitoring, where efficient video interpretation is critical.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Recognizing an action using its name: A knowledge-based approach
Chuang Gan, Yi Yang, Linchao Zhu, Deli Zhao, and Yueting Zhuang · 2016
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description, 2016
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Earlier work this paper cites.
Carla: An open urban driving simulator, 2017
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Video representation learning with deep neural networks
Linchao Zhu · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language, 2019
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh et al · 2019
Earlier work this paper cites.
Compositional attention networks with two-stream fusion for video question answering
Ting Yu, Jun Yu, Zhou Yu, and Dacheng Tao · 2020
Earlier work this paper cites.
Next-qa:next phase of question-answering to explaining temporal actions, 2021
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events
Li Xu, He Huang, and Jun Liu · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Pan Lu et al · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
et al. Minesh Mathew · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain et al · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo et al · 2021
Earlier work this paper cites.
Video question answering: Datasets, algorithms and challenges, 2022
Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li, Weihong Deng, and Tat-Seng Chua · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video, 2022
Kristen Grauman et al · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry et. al · 2022
Cited alongside, same era.
Infographicvqa
Minesh Mathew et al · 2022
Cited alongside, same era.
Advancing high-resolution video-language representation with large-scale video transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo · 2022
Cited alongside, same era.
Traffic-domain video question answering with automatic captioning, 2023
Ehsan Qasemi, Jonathan M. Francis, and Alessandro Oltramari · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik · 2023
Align and aggregate: Compositional reasoning with video alignment and answer aggregation for video question-answering, 2024
Zhaohe Liao, Jiangtong Li, Li Niu, and Liqing Zhang · 2024
Closest in time.
GPT-4o: System Card
OpenAI · 2024
Closest in time.
Dynamic spatio-temporal graph reasoning for videoqa with self-supervised event recognition
Jie Nie, Xin Wang, Runze Hou, Guohao Li, Hong Chen, and Wenwu Zhu · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team et al · 2024
Closest in time.
Video instruction tuning with synthetic data, 2024
Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li · 2024
Closest in time.
Cinepile: A long video question answering dataset and benchmark
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trecvid 2023 - a series of evaluation tracks in video understanding
G. Awad, K. Curtis, A. A. Butt, J. Fiscus, A. Godil, Y. Lee, A. Delgado, E. Godard, L. Diduch, D. Gupta, D. D. Fushman, Y. Graham, and G. Qu’enot · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu et al · 2023
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2023
Cited alongside, same era.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Cited alongside, same era.
On the hidden mystery of ocr in large multimodal models
Yuliang Liu et al · 2023
Cited alongside, same era.
Beats: Audio pre-training with acoustic tokenizers
Chen et al · 2023
Cited alongside, same era.
Ruchit Rawal, Khalid Saifullah, Ronen Basri, David Jacobs, Gowthami Somepalli, and Tom Goldstein · 2024
Closest in time.
Mist: Medical image segmentation transformer with convolutional attention mixing (cam) decoder
Md Motiur Rahman, Shiva Shokouhmand, Smriti Bhatt, and Miad Faezipour · 2024
Closest in time.
Argos vision: Advanced computer vision solutions, 2024
Argos Vision · 2024
Closest in time.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li · 2024
Closest in time.
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Closest in time.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms, 2024
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing · 2024
Closest in time.
Xiaoyi Dong et al · 2024
Closest in time.
Zheng Cai et al · 2024
Closest in time.
Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024
Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao · 2024
Closest in time.
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Chen et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li · 2024
Closest in time.