Fetching the paper…
Reading the bibliography…
The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understanding, utilizing video tokens as contextual input.
R. F. Thompson and S. A. Madigan, Memory: the key to consciousness . Princeton University Press, 2013, vol. 3
2013
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
Z. Yuan, S. Sun, L. Duan, C. Li, X. Wu, and C. Xu, “Adversarial multimodal network for movie story question answering,” IEEE Transactions on Multimedia , vol. 23, pp. 1744–1756, 2020
2020
Earlier work this paper cites.
W. Zhang, S. Tang, Y. Cao, S. Pu, F. Wu, and Y. Zhuang, “Frame augmented alternating attention network for video question answering,” IEEE Transactions on Multimedia , vol. 22, no. 4, pp. 1032–1041, 2020
2020
Earlier work this paper cites.
S. Castro, M. Azab, J. Stroud, C. Noujaim, R. Wang, J. Deng, and R. Mihalcea, “LifeQA: A real-life dataset for video question answering,” in Proceedings of the Twelfth Language Resources and Evaluation Conference , 2020, pp. 4352–4358
2020
Earlier work this paper cites.
J. Wang, B.-K. Bao, and C. Xu, “Dualvgr: A dual-visual graph reasoning unit for video question answering,” IEEE Transactions on Multimedia , vol. 24, pp. 3369–3380, 2021
2021
Earlier work this paper cites.
Z. Guo, J. Zhao, L. Jiao, X. Liu, and F. Liu, “A universal quaternion hypergraph network for multimodal video question answering,” IEEE Transactions on Multimedia , vol. 25, pp. 38–49, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 9777–9786
2021
Earlier work this paper cites.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 1686–1697
2021
Earlier work this paper cites.
J. Jiang, Z. Liu, and N. Zheng, “Livlr: A lightweight visual-linguistic reasoning framework for video question answering,” IEEE Transactions on Multimedia , vol. 25, pp. 5002–5013, 2022
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
S. Castro, N. Deng, P. Huang, M. G. Burzo, and R. Mihalcea, “In-the-wild video question answering,” in International Conference on Computational Linguistics , 2022, pp. 5613–5635
2022
Earlier work this paper cites.
L. Bärmann and A. Waibel, “Where did i leave my keys? - episodic-memory-based question answering on egocentric videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops , June 2022, pp. 1560–1568
2022
Earlier work this paper cites.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-shot video question answering via frozen bidirectional language models,” Advances in Neural Information Processing Systems , vol. 35, pp. 124–141, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Xiao, A. Yao, Z. Liu, Y. Li, W. Ji, and T.-S. Chua, “Video as conditional graph hierarchy for multi-granular question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 3, 2022, pp. 2804–2812
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Qian, R. Cui, J. Chen, P. Peng, X. Guo, and Y.-G. Jiang, “Locate before answering: Answer guided question localization for video question answering,” IEEE Transactions on Multimedia , 2023
2023
Cited alongside, same era.
F. Zhang, R. Wang, F. Zhou, Y. Luo, and J. Li, “Psam: Parameter-free spatiotemporal attention mechanism for video question answering,” IEEE Transactions on Multimedia , 2023
2023
Cited alongside, same era.
2023
Closest in time.
L. Momeni, M. Caron, A. Nagrani, A. Zisserman, and C. Schmid, “Verbs in action: Improving verb understanding in video-language models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 15 579–15 591
2023
Closest in time.
2024
Closest in time.
Y. Weng, M. Han, H. He, X. Chang, and B. Zhuang, “Longvlm: Efficient long video understanding via large language models,” in European Conference on Computer Vision . Springer, 2024, pp. 453–470
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Cheng, H. Fan, D. Lin, Y. Sun, M. Kankanhalli, and J.-H. Lim, “Keyword-aware relative spatio-temporal graph networks for video question answering,” IEEE Transactions on Multimedia , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Surís, S. Menon, and C. Vondrick, “Vipergpt: Visual inference via python execution for reasoning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 888–11 898
2023
Cited alongside, same era.
W. Dai, J. Li, D. Li, A. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” Advances in Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Yu, J. Cho, P. Yadav, and M. Bansal, “Self-chained image-language model for video localization and question answering,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhang et al. , “Moviechat: From dense token to sparse memory for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 18 221–18 232
2024
Closest in time.
Y. Li, C. Wang, and J. Jia, “Llama-vid: An image is worth 2 tokens in large language models,” in European Conference on Computer Vision , 2024
2024
Closest in time.
2024
Closest in time.
M. C. Lin and S. Yang, “Vila: Efficient video-language alignment for video question answering,” in European Conference on Computer Vision . Springer, 2024
2024
Closest in time.
H. Wang, C. Lai, Y. Sun, and W. Ge, “Weakly supervised gaussian contrastive grounding with large multimodal models for video question answering,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 5289–5298
2024
Closest in time.
R. Qian, X. Dong, P. Zhang, Y. Zang, S. Ding, D. Lin, and J. Wang, “Streaming long video understanding with large language models,” Advances in Neural Information Processing Systems , vol. 37, pp. 119 336–119 360, 2024
2024
Closest in time.
2024
Closest in time.
K. Mangalam, R. Akshulakov, and J. Malik, “Egoschema: A diagnostic benchmark for very long-form video language understanding,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
R. Liu, C. Li, Y. Ge, T. H. Li, Y. Shan, and G. Li, “Bt-adapter: Video conversation is feasible without video instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 658–13 667
2024
Closest in time.
P. Jin, R. Takanobu, W. Zhang, X. Cao, and L. Yuan, “Chat-univi: Unified visual representation empowers large language models with image and video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 700–13 710
2024
Closest in time.
2024
Closest in time.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo et al. , “Mvbench: A comprehensive multi-modal video understanding benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 195–22 206
2024
Closest in time.
S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 313–14 323
2024
Closest in time.