Fetching the paper…
Reading the bibliography…
Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing.
Sample estimate of the entropy of a random vector
Kozachenko, L. F.; and Leonenko, N. N. 1987 · 1987
Earlier work this paper cites.
k-means++: The advantages of careful seeding
Arthur, D.; Vassilvitskii, S.; et al. 2007 · 2007
Earlier work this paper cites.
A unifying view on dataset shift in classification
Moreno-Torres, J. G.; Raeder, T.; Alaiz-Rodríguez, R.; Chawla, N. V.; and Herrera, F. 2012 · 2012
Earlier work this paper cites.
Creating summaries from user videos
Gygli, M.; Grabner, H.; Riemenschneider, H.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
Fabian Caba Heilbron, B. G., Victor Escorcia; and Niebles, J. C. 2015 · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Song, Y.; Vallmitjana, J.; Stent, A.; and Jaimes, A. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Earlier work this paper cites.
Localizing Moments in Video with Temporal Language
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2018 · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Zhou, L.; Xu, C.; and Corso, J. 2018 · 2018
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019 · 2019
Earlier work this paper cites.
SODA: Story oriented dense video captioning evaluation framework
Fujita, S.; Hirao, T.; Kamigaito, H.; Okumura, M.; and Nagata, M. 2020 · 2020
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
Lei, J.; Berg, T. L.; and Bansal, M. 2021 · 2021
Earlier work this paper cites.
Queryd: A video dataset with high-quality text and audio narrations
Oncescu, A.-M.; Henriques, J. F.; Liu, Y.; Zisserman, A.; and Albanie, S. 2021 · 2021
Earlier work this paper cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Xu, H.; Ghosh, G.; Huang, P.-Y.; Okhonko, D.; Aghajanyan, A.; Metze, F.; Zettlemoyer, L.; and Feichtenhofer, C. 2021 · 2021
Earlier work this paper cites.
MERLOT: Multimodal Neural Script Knowledge Models
Zellers, R.; Lu, X.; Hessel, J.; Yu, Y.; Park, J. S.; Cao, J.; Farhadi, A.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z.; Song, Y.; Wang, J.; and Wang, L. 2022 · 2022
Cited alongside, same era.
Internvideo: General video foundation models via generative and discriminative learning
Wang, Y.; Li, K.; Li, Y.; He, Y.; Huang, B.; Zhao, Z.; Zhang, H.; Xu, J.; Liu, Y.; Wang, Z.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset
Chen, S.; Li, H.; Wang, Q.; Zhao, Z.; Sun, M.; Zhu, X.; and Liu, J. 2024 · 2024
Closest in time.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Cheng, Z.; Leng, S.; Zhang, H.; Xin, Y.; Li, X.; Chen, G.; Zhu, Y.; Zhang, W.; Luo, Z.; Zhao, D.; et al. 2024 · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Ghosal, D.; Majumder, N.; Mehrish, A.; and Poria, S. 2023 · 2023
Cited alongside, same era.
Vtimellm: Empower llm to grasp video moments
Huang, B.; Wang, X.; Chen, H.; Song, Z.; and Zhu, W. 2023 · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Cited alongside, same era.
From sparse to soft mixtures of experts
Puigcerver, J.; Riquelme, C.; Mustafa, B.; and Houlsby, N. 2023 · 2023
Cited alongside, same era.
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Ren, S.; Yao, L.; Li, S.; Sun, X.; and Hou, L. 2023 · 2023
Cited alongside, same era.
Eva-clip: Improved training techniques for clip at scale
Sun, Q.; Fang, Y.; Wu, L.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Cited alongside, same era.
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P. N.; and Hoi, S. 2024 · 2024
Closest in time.
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Guo, Y.; Liu, J.; Li, M.; Tang, X.; Chen, X.; and Zhao, B. 2024 · 2024
Closest in time.
Unleash the Potential of CLIP for Video Highlight Detection
Han, D.; Seo, S.; Park, E.; Nam, S.-U.; and Kwak, N. 2024 · 2024
Closest in time.
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
Hua, H.; Tang, Y.; Xu, C.; and Luo, J. 2024 · 2024
Closest in time.
LITA: Language Instructed Temporal-Localization Assistant
Huang, D.-A.; Liao, S.; Radhakrishnan, S.; Yin, H.; Molchanov, P.; Yu, Z.; and Kautz, J. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Qian, L.; Li, J.; Wu, Y.; Ye, Y.; Fei, H.; Chua, T.-S.; Zhuang, Y.; and Tang, S. 2024 · 2024
Closest in time.
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
Wang, Y.; Wang, Y.; Wu, P.; Liang, J.; Zhao, D.; Liu, Y.; and Zheng, Z. 2024c · 2024
Closest in time.
Number it: Temporal Grounding Videos like Flipping Manga
Wu, Y.; Hu, X.; Sun, Y.; Zhou, Y.; Zhu, W.; Rao, F.; Schiele, B.; and Yang, X. 2024 · 2024
Closest in time.
VideoPrism: A Foundational Visual Encoder for Video Understanding
Zhao, L.; Gundavarapu, N. B.; Yuan, L.; Zhou, H.; Yan, S.; Sun, J. J.; Friedman, L.; Qian, R.; Weyand, T.; Zhao, Y.; et al. 2024 · 2024
Closest in time.