Fetching the paper…
Reading the bibliography…
Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
1901
Earlier work this paper cites.
M. G. Kendall, “The treatment of ties in ranking problems,”
1945
Earlier work this paper cites.
D. Zwillinger and S. Kokoska,
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in
2005
Earlier work this paper cites.
G. Irie, T. Satou, A. Kojima, T. Yamasaki, and K. Aizawa, “Automatic trailer generation,” in
2010
Earlier work this paper cites.
S. E. F. De Avila, A. P. B. Lopes, A. da Luz Jr, and A. de Albuquerque Araújo, “Vsumm: A mechanism designed to produce static video summaries and a novel evaluation method,”
2011
Earlier work this paper cites.
D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in
2011
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,”
2012
Earlier work this paper cites.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,”
2013
Earlier work this paper cites.
M. Gygli, H. Grabner, H. Riemenschneider, and L. Van Gool, “Creating summaries from user videos,” in
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
M. Sun, A. Farhadi, and S. Seitz, “Ranking domain-specific highlights by analyzing edited videos,” in
2014
Earlier work this paper cites.
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, “Tvsum: Summarizing web videos using titles,” in
2015
Earlier work this paper cites.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” in
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in
2016
Earlier work this paper cites.
K. Zhang, W.-L. Chao, F. Sha, and K. Grauman, “Video summarization with long short-term memory,” in
2016
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in
2017
Earlier work this paper cites.
B.-C. Chen, Y.-Y. Chen, and F. Chen, “Video to text summary: Joint video summarization and captioning with recurrent neural networks,” in
2017
Earlier work this paper cites.
B. Zhao, X. Li, and X. Lu, “Hierarchical recurrent neural network for video summarization,” in
2017
Earlier work this paper cites.
——, “SGDR: Stochastic gradient descent with warm restarts,” in
2017
Earlier work this paper cites.
L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” in
2018
Earlier work this paper cites.
K. Zhou, Y. Qiao, and T. Xiang, “Deep reinforcement learning for unsupervised video summarization with diversity-representativeness reward,” in
2018
Earlier work this paper cites.
——, “Hsa-rnn: Hierarchical structure-adaptive rnn for video summarization,” in
2018
Earlier work this paper cites.
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu, “Bidirectional attentive fusion with context gating for dense video captioning,” in
2018
Earlier work this paper cites.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in
2018
Earlier work this paper cites.
B. Zhang, H. Hu, and F. Sha, “Cross-modal and hierarchical modeling of video and text,” in
2018
Earlier work this paper cites.
Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei, “Jointly localizing and describing events for dense video captioning,” in
2018
Earlier work this paper cites.
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang, “Stronger generalization bounds for deep nets via a compression approach,” in
2018
Cited alongside, same era.
S. Lal, S. Duggal, and I. Sreedevi, “Online video summarization: Predicting future to better summarize present,” in
2019
Cited alongside, same era.
Y. Yuan, H. Li, and Q. Wang, “Spatiotemporal modeling for video summarization using convolutional recurrent neural network,”
2019
Cited alongside, same era.
T.-J. Fu, S.-H. Tai, and H.-T. Chen, “Attentive and adversarial learning for video summarization,” in
2019
Cited alongside, same era.
Y. Zhang, M. Kampffmeyer, X. Zhao, and M. Tan, “Dtr-gan: Dilated temporal relational adversarial network for generic video summarization,” in
2019
Cited alongside, same era.
M. Narasimhan, A. Rohrbach, and T. Darrell, “Clip-it! language-guided video summarization,” in
2021
Later among the works it cites.
J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y. Choi, “Clipscore: A reference-free evaluation metric for image captioning,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Saquil, D. Chen, Y. He, C. Li, and Y.-L. Yang, “Multiple pairwise ranking networks for personalized video summarization,” in
2021
Later among the works it cites.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” in
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Yan, Y. Tu, X. Wang, Y. Zhang, X. Hao, Y. Zhang, and Q. Dai, “Stat: Spatial-temporal attention mechanism for video captioning,”
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Cited alongside, same era.
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in
2019
Cited alongside, same era.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in
2019
Cited alongside, same era.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in
2019
Cited alongside, same era.
H. Xue, Y. Huang, B. Liu, H. Peng, J. Fu, H. Li, and J. Luo, “Probing inter-modality: Visual parsing with self-attention for vision-language pre-training,” in
2021
Later among the works it cites.
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao, “Vinvl: Revisiting visual representations in vision-language models,” in
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in
2021
Later among the works it cites.
L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Later among the works it cites.
J. Lei, T. L. Berg, and M. Bansal, “Detecting moments and highlights in videos via natural language queries,” in
2021
Later among the works it cites.
M. Patrick, P.-Y. Huang, Y. Asano, F. Metze, A. G. Hauptmann, J. F. Henriques, and A. Vedaldi, “Support-set bottlenecks for video-text representation learning,” in
2021
Later among the works it cites.
H. Hua, X. Li, D. Dou, C. Xu, and J. Luo, “Noise stability regularization for improving BERT fine-tuning,” in
2021
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in
2022
Later among the works it cites.
C. Ju, T. Han, K. Zheng, Y. Zhang, and W. Xie, “Prompting visual-language models for efficient video understanding,” in
2022
Later among the works it cites.
B. Ni, H. Peng, M. Chen, S. Zhang, G. Meng, J. Fu, S. Xiang, and H. Ling, “Expanding language-image pretrained models for general video recognition,” in
2022
Later among the works it cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds
2022
Later among the works it cites.
J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, and L. Wang, “Git: A generative image-to-text transformer for vision and language,”
2022
Later among the works it cites.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,”
2022
Later among the works it cites.
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “VLMo: Unified vision-language pre-training with mixture-of-modality-experts,” 2022
2022
Later among the works it cites.
Y. Du, Z. Liu, J. Li, and W. X. Zhao, “A survey of vision-language pre-trained models,” in
2022
Later among the works it cites.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “Flava: A foundational language and vision alignment model,” in
2022
Later among the works it cites.
Q. Ruan, M. Ostendorff, and G. Rehm, “Histruct+: Improving extractive text summarization with hierarchical structure information,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Hu, H. Hua, Z. Yang, W. Shi, N. A. Smith, and J. Luo, “Promptcap: Prompt-guided task-aware image captioning,” in
2023
Closest in time.
H. Xu, Q. Ye, M. Yan, Y. Shi, J. Ye, Y. Xu, C. Li, B. Bi, Q. Qian, W. Wang
2023
Closest in time.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman, “Make-a-video: Text-to-video generation without text-video data,” in
2023
Closest in time.
OpenAI, “GPT-4 technical report,”
2023
Closest in time.
2023
Closest in time.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “HERO: Hierarchical encoder for Video+Language omni-representation pre-training,” in
2065
Closest in time.