Fetching the paper…
Reading the bibliography…
Video summarization aims to create short, accurate, and cohesive summaries of longer videos.
Bertscore: Evaluating text generation with bert
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
Multimodal abstractive summarization for how2 videos
Palaskar, S.; Libovickỳ, J.; Gella, S.; and Metze, F. 2019 · 1906
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
Morris, J. X.; Lifland, E.; Yoo, J. Y.; Grigsby, J.; Jin, D.; and Qi, Y. 2020 · 2005
Earlier work this paper cites.
Multi-modal summarization for video-containing documents
Fu, X.; Wang, J.; and Yang, Z. 2020 · 2009
Earlier work this paper cites.
Support-set bottlenecks for video-text representation learning
Patrick, M.; Huang, P.-Y.; Asano, Y.; Metze, F.; Hauptmann, A.; Henriques, J.; and Vedaldi, A. 2020 · 2010
Earlier work this paper cites.
Collecting Highly Parallel Data for Paraphrase Evaluation
Chen, D. L.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method
de Avila, S. E. F.; Lopes, A. P. B.; da Luz, A.; and de Albuquerque Araújo, A. 2011 · 2011
Earlier work this paper cites.
VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method
De Avila, S. E. F.; Lopes, A. P. B.; da Luz Jr, A.; and de Albuquerque Araújo, A. 2011 · 2011
Earlier work this paper cites.
Multimodal pretraining for dense video captioning
Huang, G.; Pang, B.; Zhu, Z.; Rivera, C.; and Soricut, R. 2020 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Das, P.; Xu, C.; Doell, R. F.; and Corso, J. J. 2013 · 2013
Earlier work this paper cites.
Creating summaries from user videos
Gygli, M.; Grabner, H.; Riemenschneider, H.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Category-specific video summarization
Potapov, D.; Douze, M.; Harchaoui, Z.; and Schmid, C. 2014 · 2014
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Song, Y.; Vallmitjana, J.; Stent, A.; and Jaimes, A. 2015 · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Zhang, K.; Chao, W.-L.; Sha, F.; and Grauman, K. 2016 · 2016
Earlier work this paper cites.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Earlier work this paper cites.
Query-focused video summarization: Dataset, evaluation, and a memory network based approach
Sharghi, A.; Laurel, J. S.; and Gong, B. 2017 · 2017
Earlier work this paper cites.
Contextually customized video summaries via natural language
Choi, J.; Oh, T.-H.; and Kweon, I. S. 2018 · 2018
Earlier work this paper cites.
Summarizing videos with attention
Fajtl, J.; Sokeh, H. S.; Argyriou, V.; Monekosso, D.; and Remagnino, P. 2019 · 2018
Earlier work this paper cites.
Extractive video summarizer with memory augmented neural networks
Feng, L.; Li, Z.; Kuang, Z.; and Zhang, W. 2018 · 2018
Earlier work this paper cites.
Jointly localizing and describing events for dense video captioning
Li, Y.; Yao, T.; Pan, Y.; Chao, H.; and Mei, T. 2018 · 2018
Earlier work this paper cites.
Retrospective encoders for video summarization
Zhang, K.; Grauman, K.; and Sha, F. 2018 · 2018
Cited alongside, same era.
Deep reinforcement learning for unsupervised video summarization with diversity-representativeness reward
Zhou, K.; Qiao, Y.; and Xiang, T. 2018 · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Zhou, L.; Xu, C.; and Corso, J. 2018 · 2018
Cited alongside, same era.
End-to-end dense video captioning with masked transformer
Zhou, L.; Zhou, Y.; Corso, J. J.; Socher, R.; and Xiong, C. 2018 · 2018
Cited alongside, same era.
Attentive and Adversarial Learning for Video Summarization
Fu, T.-J.; Tai, S.-H.; and Chen, H.-T. 2019 · 2019
Cited alongside, same era.
Video summarization with attention-based encoder–decoder networks
Ji, Z.; Xiong, K.; Pang, Y.; and Li, X. 2019 · 2019
Joint Video Summarization and Moment Localization by Cross-Task Sample Transfer
Jiang, H.; and Mu, Y. 2022 · 2022
Later among the works it cites.
Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency
Narasimhan, M.; Nagrani, A.; Sun, C.; Rubinstein, M.; Darrell, T.; Rohrbach, A.; and Schmid, C. 2022 · 2022
Later among the works it cites.
Multi-modal segment assemblage network for ad video editing with importance-coherence reward
Tang, Y.; Xu, S.; Wang, T.; Lin, Q.; Lu, Q.; and Zheng, F. 2022 · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J.; Tay, Y.; Bommasani, R.; Raffel, C.; Zoph, B.; Borgeaud, S.; Yogatama, D.; Bosma, M.; Zhou, D.; Metzler, D.; et al. 2022 · 2022
Later among the works it cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Susinet: See, understand and summarize it
Koutras, P.; and Maragos, P. 2019 · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Rethinking the evaluation of video summaries
Otani, M.; Nakashima, Y.; Rahtu, E.; and Heikkila, J. 2019 · 2019
Cited alongside, same era.
Stacked memory network for video summarization
Wang, J.; Wang, W.; Wang, Z.; Wang, L.; Feng, D.; and Tan, T. 2019 · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z.; Xu, D.; Yu, J.; Yu, T.; Zhao, Z.; Zhuang, Y.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Classification of Important Segments in Educational Videos using Multimodal Features
Ghauri, J. A.; Hakimov, S.; and Ewerth, R. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Han, M.; Chang, X.; Wang, H.; and Yang, L. 2023 · 2023
Later among the works it cites.
Improving Pretrained Language Model Fine-Tuning With Noise Stability Regularization
Hua, H.; Li, X.; Dou, D.; Xu, C.-Z.; and Luo, J. 2023 · 2023
Later among the works it cites.
Vtimellm: Empower llm to grasp video moments
Huang, B.; Wang, X.; Chen, H.; Song, Z.; and Zhu, W. 2023 · 2023
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Lyu, C.; Wu, M.; Wang, L.; Huang, X.; Liu, B.; Du, Z.; Shi, S.; and Tu, Z. 2023 · 2023
Later among the works it cites.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Later among the works it cites.
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Ren, S.; Yao, L.; Li, S.; Sun, X.; and Hou, L. 2023 · 2023
Later among the works it cites.
Audio-Visual LLM for Video Understanding
Shu, F.; Zhang, L.; Jiang, H.; and Xie, C. 2023 · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Su, Y.; Lan, T.; Li, H.; Xu, J.; Wang, Y.; and Cai, D. 2023 · 2023
Later among the works it cites.
Video understanding with large language models: A survey
Tang, Y.; Bi, J.; Xu, S.; Song, L.; Liang, S.; Wang, T.; Zhang, D.; An, J.; Lin, J.; Zhu, R.; et al. 2023 · 2023
Later among the works it cites.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Yang, A.; Nagrani, A.; Seo, P. H.; Miech, A.; Pont-Tuset, J.; Laptev, I.; Sivic, J.; and Schmid, C. 2023 · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
Diffusum: Generation enhanced extractive summarization with diffusion
Zhang, H.; Liu, X.; and Zhang, J. 2023 · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Zhang, R.; Han, J.; Liu, C.; Gao, P.; Zhou, A.; Hu, X.; Yan, S.; Lu, P.; Li, H.; and Qiao, Y. 2023 · 2023
Later among the works it cites.
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
Hua, H.; Shi, J.; Kafle, K.; Jenni, S.; Zhang, D.; Collomosse, J.; Cohen, S.; and Luo, J. 2024 · 2024
Closest in time.
LEGO: Language Enhanced Multi-modal Grounding Model
Li, Z.; Xu, Q.; Zhang, D.; Song, H.; Cai, Y.; Qi, Q.; Zhou, R.; Pan, J.; Li, Z.; Vu, V. T.; et al. 2024 · 2024
Closest in time.
Tang, Y.; Shimada, D.; Bi, J.; and Xu, C. 2024 · 2024
Closest in time.
Vidchapters-7m: Video chapters at scale
Yang, A.; Nagrani, A.; Laptev, I.; Sivic, J.; and Schmid, C. 2024 · 2024
Closest in time.
PromptFix: You Prompt and We Fix the Photo
Yu, Y.; Zeng, Z.; Hua, H.; Fu, J.; and Luo, J. 2024 · 2024
Closest in time.