Fetching the paper…
Reading the bibliography…
The core challenge in video understanding lies in perceiving dynamic content changes over time.
E. Tulving, “Episodic and semantic memory,” Organization of memory , vol. 1, no. 381-403, p. 1, 1972
1972
Earlier work this paper cites.
T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the h. 264/avc video coding standard,” IEEE Transactions on circuits and systems for video technology , vol. 13, no. 7, pp. 560–576, 2003
2003
Earlier work this paper cites.
M. Sun, A. Farhadi, and S. Seitz, “Ranking domain-specific highlights by analyzing edited videos,” in European Conference on Computer Vision (ECCV) , 2014
2014
Earlier work this paper cites.
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, “Tvsum: Summarizing web videos using titles,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2015
2015
Earlier work this paper cites.
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2015
2015
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in IEEE/CVF International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles, “Dense-captioning events in videos,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Earlier work this paper cites.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in AAAI Conference on Artificial Intelligence (AAAI) , 2018
2018
Earlier work this paper cites.
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu, “Bidirectional attentive fusion with context gating for dense video captioning,” in IEEE conference on computer vision and pattern recognition (CVPR) , 2018
2018
Earlier work this paper cites.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in IEEE conference on computer vision and pattern recognition (CVPR) , 2018
2018
Earlier work this paper cites.
J. Lei, L. Yu, M. Bansal, and T. Berg, “TVQA: Localized, compositional video question answering,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2018
2018
Earlier work this paper cites.
L. Zhou, Y. Zhou, J. Corso, R. Socher, and C. Xiong, “End-to-end dense video cap- tioning with masked transformer,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Earlier work this paper cites.
Y. Guo, J. Liu, M. Li, D. Chen, X. Tang, D. Sui, Q. Liu, X. Chen, and K. Zhao, “Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,” in AAAI Conference on Artificial Intelligence (AAAI) , 2019
2019
Earlier work this paper cites.
H. Zhang, A. Sun, W. Jing, and J. T. Zhou, “Span-based localizing network for natural language video localization,” in Annual Meeting of the Association for Computational Linguistics (ACL) , 2020
2020
Earlier work this paper cites.
J. Mun, M. Cho, and B. Han, “Local-global video-text interactions for temporal grounding,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Earlier work this paper cites.
V. Iashin and E. Rahtu, “A better use of audio-visual cues: Dense video captioning with bi-modal transformer,” in The British Machine Vision Conference (BMVC) , 2020
2020
Earlier work this paper cites.
J. Lei, L. Yu, T. Berg, and M. Bansal, “TVQA+: Spatio-temporal grounding for video question answering,” in Annual Meeting of the Association for Computational Linguistics (ACL) , 2020
2020
Earlier work this paper cites.
J. Lei, T. L. Berg, and M. Bansal, “Detecting moments and highlights in videos via natural language queries,” in Advances in Neural Information Processing Systems (NeurIPS) , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning (ICML) , 2021
2021
Earlier work this paper cites.
S. Zhang, H. Peng, J. Fu, Y. Lu, and J. Luo, “Multi-scale 2d temporal adjacency networks for moment localization with natural language,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 12, pp. 9073–9087, 2021
2021
Earlier work this paper cites.
T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo, “End-to-end dense video captioning with parallel decoding,” in IEEE/CVF international conference on computer vision (CVPR) , 2021
2021
Earlier work this paper cites.
C. Deng, S. Chen, D. Chen, Y. He, and Q. Wu, “Sketch, ground, and refine: Top-down dense video captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Earlier work this paper cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021
2021
Earlier work this paper cites.
Y. Liu, S. Li, Y. Wu, C.-W. Chen, Y. Shan, and X. Qie, “Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations (ICLR) , 2022
2022
Cited alongside, same era.
Y. Li, X. Wang, J. Xiao, W. Ji, and T.-S. Chua, “Invariant grounding for video question answering,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2022
Cited alongside, same era.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Zero-shot video question answering via frozen bidirectional language models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning (ICML) , 2023
2023
Cited alongside, same era.
X. Wang, Y. Zhang, and S. Yeung-Levy, “Videoagent: Long-form video understanding with large language model as agent,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
D.-A. Huang, S. Liao, S. Radhakrishnan, H. Yin, P. Molchanov, Z. Yu, and J. Kautz, “Lita: Language instructed temporal-localization assistant,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
L. Qian, J. Li, Y. Wu, Y. Yaobo, F. Hao, T.-S. Chua, Y. Zhuang, and T. Siliang, “Momentor: Advancing video large language model with fine-grained temporal reasoning,” in Proceedings of the International Conference on Machine Learning (ICML) , 2024
2024
Later among the works it cites.
J. Xiao, A. Yao, Y. Li, and T.-S. Chua, “Can i trust your answer? visually grounded video question answering,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2023
2023
Cited alongside, same era.
H. Chen, X. Wang, H. Chen, Z. Song, J. Jia, and W. Zhu, “Grounding-prompter: Prompting llm with multimodal information for temporal sentence grounding in long videos,” in arXiv , 2023
2023
Cited alongside, same era.
S. Yu, J. Cho, P. Yadav, and M. Bansal, “Self-chained image-language model for video localization and question answering,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” in arXiv , 2023
2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” in arXiv , 2023
2023
Cited alongside, same era.
W. Moon, S. Hyun, S. Park, D. Park, and J.-P. Heo, “Query-dependent video representation for moment retrieval and highlight detection,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Cited alongside, same era.
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid, “Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Cited alongside, same era.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo, and Z. Fan, “Qwen2 technical report,” in arXiv , 2024
2024
Later among the works it cites.
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
Z. Ren, Z. Huang, Y. Wei, Y. Zhao, D. Fu, J. Feng, and X. Jin, “Pixellm: Pixel reasoning with large multimodal model,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
H. Rasheed, M. Maaz, S. Shaji, A. Shaker, S. Khan, H. Cholakkal, R. M. Anwer, E. Xing, M.-H. Yang, and F. S. Khan, “Glamm: Pixel grounding large multimodal model,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
K. Ataallah, X. Shen, E. Abdelrahman, E. Sleiman, M. Zhuge, J. Ding, D. Zhu, J. Schmidhuber, and M. Elhoseiny, “Goldfish: Vision-language understanding of arbitrarily long videos,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
D. Han, X. Cheng, N. Guo, X. Ye, B. Rainer, and P. Priller, “Momentum cross-modal contrastive learning for video moment retrieval,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 7, pp. 5977–5994, 2024
2024
Later among the works it cites.
X. Jiang, X. Xu, J. Zhang, F. Shen, Z. Cao, and H. T. Shen, “Sdn: Semantic decoupling network for temporal language grounding,” IEEE Transactions on Neural Networks and Learning Systems , vol. 35, no. 5, pp. 6598–6612, 2024
2024
Later among the works it cites.
X. Jiang, L. Zhu, X. Xu, F. Shen, Y. Yang, and H. T. Shen, “Query as supervision: Towards low-cost and robust video moment and highlight retrieval,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2024
2024
Later among the works it cites.
C. Deng, S. Chen, D. Chen, Y. He, and Q. Wu, “Grounded question-answering in long egocentric videos,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, P. Luo, z. Liu, Y. Wang, L. Wang, and Y. Qiao, “Internvid: A large-scale video-text dataset for multimodal understanding and generation,” in International Conference on Learning Representations (ICLR) , 2024
2024
Later among the works it cites.
Y. Wang, X. Meng, J. Liang, Y. Wang, Q. Liu, and D. Zhao, “Hawkeye: Training video-text llms for grounding text in videos,” in arXiv , 2024
2024
Later among the works it cites.
R. Qian, X. Dong, P. Zhang, Y. Zang, S. Ding, D. Lin, and J. Wang, “Streaming long video understanding with large language models,” in Advances in Neural Information Processing Systems (NeurIPS) , 2024
2024
Later among the works it cites.
R. Liu, C. Li, Y. Ge, Y. Shan, T. H. Li, and G. Li, “Bt-adapter: Video conversation is feasible without video instruction tuning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Later among the works it cites.
Y. Li, C. Wang, and J. Jia, “Llama-vid: An image is worth 2 tokens in large language models,” in European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
Q. Ye, Z. Yu, R. Shao, X. Xie, and C. Xiaochun, “Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios,” in European Conference on Computer Vision (ECCV) , 2024
2024
Later among the works it cites.
X. Chen, W. Xu, S. Kan, L. Zhang, Y. Jin, Y. Cen, and Y. Li, “Vision-semantics-label: A new two-step paradigm for action recognition with large language model,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2025
2025
Closest in time.
Z. Wang, S. Yu, E. Stengel-Eskin, J. Yoon, F. Cheng, G. Bertasius, and M. Bansal, “Videotree: Adaptive tree-based video representation for llm reasoning on long videos,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2025
2025
Closest in time.
J. Liu, Z. He, W. Nie, Z. Zhang, and Y. Su, “What and where: Semantic grasping and contextual scanning for moment retrieval and highlight detection,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2025
2025
Closest in time.
L. Zhang, Y. Liu, Z. Zhang, M. Aghaei, Y. Hu, H. Gu, M. A. Alomrani, D. G. A. Bravo, R. Karimi, A. Hamidizadeh, H. Xu, G. Huang, Z. Zhang, T. Cao, W. Qiu, X. Quan, J. Hao, Y. Zhuang, and Y. Zhang, “Mem2ego: Empowering vision-language models with global-to-ego memory for long-horizon embodied navigation,” in arXiv , 2025
2025
Closest in time.