Fetching the paper…
Reading the bibliography…
Video Question Answering (VideoQA) aims to answer natural language questions based on the information observed in videos.
Tall: Temporal activity localization via language query. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in Neural Information Processing Systems
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion. In Proceedings of the ACM International Conference on Multimedia
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia. 2018 · 2018
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019 · 2019
Earlier work this paper cites.
Learning with differentiable pertubed optimizers. In Advances in Neural Information Processing Systems
Quentin Berthet, Mathieu Blondel, Olivier Teboul, Marco Cuturi, Jean-Philippe Vert, and Francis Bach. 2020 · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Earlier work this paper cites.
Reasoning with heterogeneous graph alignment for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence
Pin Jiang and Yahong Han. 2020 · 2020
Earlier work this paper cites.
Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries. In Advances in Neural Information Processing Systems
Jie Lei, Tamara L Berg, and Mohit Bansal. 2021 · 2021
Earlier work this paper cites.
Bridge to answer: Structure-aware graph interaction network for video question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. 2021 · 2021
Cited alongside, same era.
Progressive graph attention network for video question answering. In Proceedings of the ACM International Conference on Multimedia . 2871–2879
Liang Peng, Shuangji Yang, Yi Bin, and Guoqing Wang. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. 2021 · 2021
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023 · 2023
Later among the works it cites.
MIST: Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question Answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou. 2023 · 2023
Later among the works it cites.
Visual instruction tuning. In Advances in Neural Information Processing Systems
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Revisiting the" video" in video-language understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. 2022 · 2022
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Cited alongside, same era.
Internvideo: General video foundation models via generative and discriminative learning
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Cited alongside, same era.
Video-text modeling with zero-shot transfer from contrastive captioners
Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, and Jiahui Yu. 2022 · 2022
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models. In Advances in Neural Information Processing Systems
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022 · 2022
Cited alongside, same era.
Valor: Vision-audio-language omni-perception pretraining model and dataset
Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu. 2023 · 2023
Cited alongside, same era.
All in one: Exploring unified video-language pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al · 2023
Later among the works it cites.
Visual causal scene refinement for video question answering. In Proceedings of the ACM International Conference on Multimedia . 377–386
Yushen Wei, Yang Liu, Hong Yan, Guanbin Li, and Liang Lin. 2023 · 2023
Later among the works it cites.
Can I Trust Your Answer? Visually Grounded Video Question Answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat Seng Chua. 2023 · 2023
Later among the works it cites.
mplug-2: A modularized multi-modal foundation model across text, image and video. In International Conference on Machine Learning
Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, et al · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al · 2023
Later among the works it cites.
Self-Chained Image-Language Model for Video Localization and Question Answering. In Advances in Neural Information Processing Systems
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Temporal sentence grounding in videos: A survey and future directions
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model. In International Conference on Learning Representations
Sihan Chen, Xingjian He, Handong Li, Xiaojie Jin, Jiashi Feng, and Jing Liu. 2024 · 2024
Closest in time.