Fetching the paper…
Reading the bibliography…
Video Large Language Models (Video-LLMs) are flourishing and has advanced many video-language tasks.
Agrawal A, Batra D, Parikh D (2016) Analyzing the behavior of visual question answering models. In: Conference on Empirical Methods in Natural Language Processing (EMNLP), pp 1955–1960
1960
Earlier work this paper cites.
Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural computation 9(8):1735–1780
1997
Earlier work this paper cites.
Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Zitnick CL, Parikh D (2015) Vqa: Visual question answering. In: Proceedings of the IEEE international conference on computer vision (ICCV), pp 2425–2433
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 770–778
2016
Earlier work this paper cites.
Goyal Y, Khot T, Summers-Stay D, Batra D, Parikh D (2017) Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 6904–6913
2017
Earlier work this paper cites.
Jang Y, Song Y, Yu Y, Kim Y, Kim G (2017) Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 2758–2766
2017
Earlier work this paper cites.
Xu D, Zhao Z, Xiao J, Wu F, Zhang H, He X, Zhuang Y (2017) Video question answering via gradually refined attention over appearance and motion. In: Proceedings of the 25th ACM international conference on Multimedia, pp 1645–1653
2017
Earlier work this paper cites.
Devlin J, Chang MW, Lee K, Toutanova K (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:181004805
2018
Earlier work this paper cites.
Jang Y, Song Y, Kim CD, Yu Y, Kim Y, Kim G (2019) Video question answering with spatio-temporal reasoning. International Journal of Computer Vision (IJCV) 127:1385–1412
2019
Earlier work this paper cites.
Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:190711692
2019
Earlier work this paper cites.
Shah M, Chen X, Rohrbach M, Parikh D (2019) Cycle-consistency for robust visual question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 6649–6658
2019
Earlier work this paper cites.
Sun C, Myers A, Vondrick C, Murphy K, Schmid C (2019) Videobert: A joint model for video and language representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp 7464–7473
2019
Earlier work this paper cites.
Yu Z, Xu D, Yu J, Yu T, Zhao Z, Zhuang Y, Tao D (2019) Activitynet-qa: A dataset for understanding complex web videos via question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol 33, pp 9127–9134
2019
Earlier work this paper cites.
He P, Liu X, Gao J, Chen W (2020) Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:200603654
2020
Earlier work this paper cites.
Le TM, Le V, Venkatesh S, Tran T (2020) Hierarchical conditional relation networks for video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 9972–9981
2020
Earlier work this paper cites.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al. (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: International Conference on Learning Representations (ICLR)
2021
Earlier work this paper cites.
Fu TJ, Li L, Gan Z, Lin K, Wang WY, Wang L, Liu Z (2021) Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:211112681
2021
Earlier work this paper cites.
Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W (2021) Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:210609685
2021
Earlier work this paper cites.
Ilharco G, Wortsman M, Wightman R, Gordon C, Carlini N, Taori R, Dave A, Shankar V, Namkoong H, Miller J, Hajishirzi H, Farhadi A, Schmidt L (2021) Openclip. DOI 10.5281/zenodo.5143773
2021
Earlier work this paper cites.
Kervadec C, Antipov G, Baccouche M, Wolf C (2021) Roses are red, violets are blue… but should vqa expect them to? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 2776–2785
2021
Earlier work this paper cites.
Lei J, Li L, Zhou L, Gan Z, Berg TL, Bansal M, Liu J (2021) Less is more: Clipbert for video-and-language learning via sparse sampling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 7331–7341
2021
Earlier work this paper cites.
Niu Y, Tang K, Zhang H, Lu Z, Hua XS, Wen JR (2021) Counterfactual vqa: A cause-effect look at language bias. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 12700–12710
2021
Earlier work this paper cites.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al. (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML), PMLR, pp 8748–8763
2021
Earlier work this paper cites.
Seo PH, Nagrani A, Schmid C (2021) Look before you speak: Visually contextualized utterances. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 16877–16887
2021
Earlier work this paper cites.
Xiao J, Shang X, Yao A, Chua TS (2021) Next-qa: Next phase of question-answering to explaining temporal actions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 9777–9786
2021
Earlier work this paper cites.
Yang A, Miech A, Sivic J, Laptev I, Schmid C (2021) Just ask: Learning to answer questions from millions of narrated videos. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp 1686–1697
2021
Cited alongside, same era.
Alayrac JB, Donahue J, Luc P, Miech A, Barr I, Hasson Y, Lenc K, Mensch A, Millican K, Reynolds M, et al. (2022) Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS) 35:23716–23736
2022
Cited alongside, same era.
Buch S, Eyzaguirre C, Gaidon A, Wu J, Fei-Fei L, Niebles JC (2022) Revisiting the ”video” in video-language understanding. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pp 2917–2927
2022
Cited alongside, same era.
Chung HW, Hou L, Longpre S, Zoph B, Tay Y, Fedus W, Li Y, Wang X, Dehghani M, Brahma S, et al. (2022) Scaling instruction-finetuned language models. arXiv preprint arXiv:221011416
2022
Pătrăucean V, Smaira L, Gupta A, Continente AR, Markeeva L, Banarse D, Koppula S, Heyward J, Malinowski M, Yang Y, et al. (2023) Perception test: A diagnostic benchmark for multimodal video models. In: The 37th Conference on Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks
2023
Later among the works it cites.
Sun Q, Fang Y, Wu L, Wang X, Cao Y (2023) Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:230315389
2023
Later among the works it cites.
Surís D, Menon S, Vondrick C (2023) Vipergpt: Visual inference via python execution for reasoning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 11888–11898
2023
Later among the works it cites.
Tang Y, Bi J, Xu S, Song L, Liang S, Wang T, Zhang D, An J, Lin J, Zhu R, et al. (2023) Video understanding with large language models: A survey. arXiv preprint arXiv:231217432
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Datta S, Dharur S, Cartillier V, Desai R, Khanna M, Batra D, Parikh D (2022) Episodic memory question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 19119–19128
2022
Cited alongside, same era.
Grauman K, Westbury A, Byrne E, Chavis Z, Furnari A, Girdhar R, Hamburger J, Jiang H, Liu M, Liu X, et al. (2022) Ego4d: Around the world in 3,000 hours of egocentric video. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18995–19012
2022
Cited alongside, same era.
Kojima T, Gu SS, Reid M, Matsuo Y, Iwasawa Y (2022) Large language models are zero-shot reasoners. Advances in neural information processing systems (NeurIPS) 35:22199–22213
2022
Cited alongside, same era.
Li Y, Wang X, Xiao J, Ji W, Chua TS (2022) Invariant grounding for video question answering. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 2928–2937
2022
Cited alongside, same era.
Liu Z, Ning J, Cao Y, Wei Y, Zhang Z, Lin S, Hu H (2022) Video swin transformer. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pp 3202–3211
2022
Cited alongside, same era.
Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D, et al. (2022) Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems (NeurIPS) 35:24824–24837
2022
Cited alongside, same era.
Yang A, Miech A, Sivic J, Laptev I, Schmid C (2022) Zero-shot video question answering via frozen bidirectional language models. Advances in Neural Information Processing Systems (NeurIPS) 35:124–141
2022
Cited alongside, same era.
Zhong Y, Xiao J, Ji W, Li Y, Deng W, Chua TS (2022) Video question answering: Datasets, algorithms and challenges. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp 6439–6455
2022
Cited alongside, same era.
Team G, Anil R, Borgeaud S, Wu Y, Alayrac JB, Yu J, Soricut R, Schalkwyk J, Dai AM, Hauth A, et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:231211805
2023
Later among the works it cites.
Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, Rozière B, Goyal N, Hambro E, Azhar F, et al. (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:230213971
2023
Later among the works it cites.
Xiao J, Zhou P, Yao A, Li Y, Hong R, Yan S, Chua TS (2023) Contrastive video question answering via video graph transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI) 45(11):13265–13280, DOI 10.1109/TPAMI.2023.3292266
2023
Later among the works it cites.
Yu S, Cho J, Yadav P, Bansal M (2023) Self-chained image-language model for video localization and question answering. In: The 37th Conference on Neural Information Processing Systems (NeurIPS)
2023
Later among the works it cites.
Zeng A, Attarian M, Choromanski KM, Wong A, Welker S, Tombari F, Purohit A, Ryoo MS, Sindhwani V, Lee VV Johnny, Florence P (2023) Socratic models: Composing zero-shot multimodal reasoning with language. In: The Eleventh International Conference on Learning Representations (ICLR)
2023
Later among the works it cites.
Zhang H, Li X, Bing L (2023b) Video-llama: An instruction-tuned audio-visual language model for video understanding. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp 543–553
2023
Later among the works it cites.
Zhang X, Zhang F, Xu C (2024c) Next-ood: Overcoming dual multiple-choice vqa biases. IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI) 46(4):1913–1931, DOI 10.1109/TPAMI.2023.3269429
2023
Later among the works it cites.
Zhao Y, Misra I, Krähenbühl P, Girdhar R (2023) Learning video representations from large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 6586–6597
2023
Later among the works it cites.
Zhu B, Lin B, Ning M, Yan Y, Cui J, Wang H, Pang Y, Jiang W, Zhang J, Li Z, et al. (2023) Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment. arXiv preprint arXiv:231001852
2023
Later among the works it cites.
Bai Z, Wang P, Xiao T, He T, Han Z, Zhang Z, Shou MZ (2024) Hallucination of multimodal large language models: A survey. arXiv preprint arXiv:240418930
2024
Closest in time.
Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, Mathur A, Schelten A, Yang A, Fan A, et al. (2024) The llama 3 herd of models. arXiv preprint arXiv:240721783
2024
Closest in time.
Fan Y, Ma X, Wu R, Du Y, Li J, Gao Z, Li Q (2024) Videoagent: A memory-augmented multimodal agent for video understanding. European Conference on Computer Vision (ECCV)
2024
Closest in time.
Fei H, Wu S, Ji W, Zhang H, Zhang M, Lee ML, Hsu W (2024) Video-of-thought: Step-by-step video reasoning from perception to cognition. In: Forty-first International Conference on Machine Learning (ICML)
2024
Closest in time.
Fu C, Dai Y, Luo Y, Li L, Ren S, Zhang R, Wang Z, Zhou C, Shen Y, Zhang M, et al. (2024) Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:240521075
2024
Closest in time.
Kim W, Choi C, Lee W, Rhee W (2024) An image grid can be worth a video: Zero-shot video question answering using a vlm. arXiv preprint arXiv:240318406
2024
Closest in time.
Majumdar A, Ajay A, Zhang X, Putta P, Yenamandra S, Henaff M, Silwal S, Mcvay P, Maksymets O, Arnaud S, et al. (2024) Openeqa: Embodied question answering in the era of foundation models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 16488–16498
2024
Closest in time.
Min J, Buch S, Nagrani A, Cho M, Schmid C (2024) Morevqa: Exploring modular reasoning models for video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 13235–13245
2024
Closest in time.
Shang C, You A, Subramanian S, Darrell T, Herzig R (2024) Traveler: A multi-lmm agent framework for video question-answering. arXiv preprint arXiv:240401476
2024
Closest in time.
Xiao J, Yao A, Li Y, Chua TS (2024) Can i trust your answer? visually grounded video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 13204–13214
2024
Closest in time.
Zhang Y, Li B, Liu h, Lee Yj, Gui L, Fu D, Feng J, Liu Z, Li C (2024e) Llava-next: A strong zero-shot video understanding model. URL https://llava-vl.github.io/blog/2024-04-30-llava-next-video/
2024
Closest in time.
Li L, Chen YC, Cheng Y, Gan Z, Yu L, Liu J (2020) Hero: Hierarchical encoder for video+ language omni-representation pre-training. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp 2046–2065
2065
Closest in time.