Fetching the paper…
Reading the bibliography…
Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments paired with text.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Video captioning with transferred semantic attributes
Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei · 2017
Earlier work this paper cites.
Weakly supervised dense event captioning in videos
Xuguang Duan, Wenbing Huang, Chuang Gan, Jingdong Wang, Wenwu Zhu, and Junzhou Huang · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso · 2018
Earlier work this paper cites.
Localizing natural language in videos
Jingyuan Chen, Lin Ma, Xinpeng Chen, Zequn Jie, and Jiebo Luo · 2019
Earlier work this paper cites.
Wslln: Weakly supervised natural language localization networks
Mingfei Gao, Larry S Davis, Richard Socher, and Caiming Xiong · 2019
Earlier work this paper cites.
Debug: A dense bottom-up grounding approach for natural language video localization
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao · 2019
Earlier work this paper cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K Roy-Chowdhury · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Watch, listen and tell: Multi-modal weakly supervised dense event captioning
Tanzila Rahman, Bicheng Xu, and Leonid Sigal · 2019
Earlier work this paper cites.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Earlier work this paper cites.
Soda: Story oriented dense video captioning evaluation framework
Soichiro Fujita, Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, and Masaaki Nagata · 2020
Earlier work this paper cites.
Vlanet: Video-language alignment network for weakly-supervised video moment retrieval
Minuk Ma, Sunjae Yoon, Junyeong Kim, Youngjoon Lee, Sunghun Kang, and Chang D Yoo · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2020
Cited alongside, same era.
Proposal-free temporal moment localization of a natural-language query in video using guided attention
Cristian Rodriguez, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould · 2020
Cited alongside, same era.
Event-centric hierarchical representation for dense video captioning
Teng Wang, Huicheng Zheng, Mingjing Yu, Qian Tian, and Haifeng Hu · 2020
Cited alongside, same era.
Dense regression network for video grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan · 2020
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Weakly-supervised moment retrieval network for video corpus moment retrieval
Sunjae Yoon, Dahyun Kim, Ji Woo Hong, Junyeong Kim, Kookhoi Kim, and Chang D Yoo · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Pseudo-q: Generating pseudo language queries for visual grounding
Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, and Gao Huang · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Swinbert: End-to-end transformers with sparse attention for video captioning
Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning
Shaoxiang Chen and Yu-Gang Jiang · 2021
Cited alongside, same era.
Sketch, ground, and refine: Top-down dense video captioning
Chaorui Deng, Shizhe Chen, Da Chen, Yuan He, and Qi Wu · 2021
Cited alongside, same era.
Magma–multimodal augmentation of generative models through adapter-based finetuning
Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, and Anette Frank · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Zero-shot natural language video localization
Jinwoo Nam, Daechul Ahn, Dongyeop Kang, Seong Jong Ha, and Jonghyun Choi · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Linearly mapping from image to text space
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick · 2022
Later among the works it cites.
Text-based temporal localization of novel events
Sudipta Paul, Niluthpol Chowdhury Mithun, and Amit K Roy-Chowdhury · 2022
Later among the works it cites.
Prompt-based zero-shot video moment retrieval
Guolong Wang, Xun Wu, Zhaoyuan Liu, and Junchi Yan · 2022
Later among the works it cites.
Unifying event detection and captioning as sequence generation via pre-training
Qi Zhang, Yuqing Song, and Qin Jin · 2022
Later among the works it cites.
End-to-end dense video captioning as sequence generation
Wanrong Zhu, Bo Pang, Ashish Thapliyal, William Yang Wang, and Radu Soricut · 2022
Later among the works it cites.
Language-free training for zero-shot video grounding
Dahye Kim, Jungin Park, Jiyoung Lee, Seongheon Park, and Kwanghoon Sohn · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid · 2023
Closest in time.