Fetching the paper…
Reading the bibliography…
There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video.
An event-related potential study of explicit memory on tests of cued recall and recognition
Ken Allan and MD Rugg · 1997
Earlier work this paper cites.
Neural correlates of memory retrieval during recognition memory and cued recall
Michael D Rugg, Paul C Fletcher, Kevin Allan, Chris D Frith, RSJ Frackowiak, and Raymond J Dolan · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Vladimir Iashin and Esa Rahtu · 2005
Earlier work this paper cites.
Dense-captioning events in videos: Sysu submission to activitynet challenge 2020
Teng Wang, Huicheng Zheng, and Mingjing Yu · 2006
Earlier work this paper cites.
Translating video content to natural language descriptions
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Earlier work this paper cites.
Video captioning with guidance of multimodal latent topics
Shizhe Chen, Jia Chen, Qin Jin, and Alexander Hauptmann · 2017
Earlier work this paper cites.
Video captioning with attention-based lstm and semantic consistency
Lianli Gao, Zhao Guo, Hanwang Zhang, Xing Xu, and Heng Tao Shen · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Weakly supervised dense video captioning
Zhiqiang Shen, Jianguo Li, Zhou Su, Minjun Li, Yurong Chen, Yu-Gang Jiang, and Xiangyang Xue · 2017
Earlier work this paper cites.
Jointly localizing and describing events for dense video captioning
Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei · 2018
Earlier work this paper cites.
Hierarchical context encoding for events captioning in videos
Dali Yang and Chun Yuan · 2018
Earlier work this paper cites.
Streamlined dense video captioning
Jonghwan Mun, Linjie Yang, Zhou Ren, Ning Xu, and Bohyung Han · 2019
Cited alongside, same era.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai · 2019
Cited alongside, same era.
Sports video captioning via attentive motion representation and group relationship modeling
Mengshi Qi, Yunhong Wang, Annan Li, and Jiebo Luo · 2019
Cited alongside, same era.
Watch, listen and tell: Multi-modal weakly supervised dense event captioning
Tanzila Rahman, Bicheng Xu, and Leonid Sigal · 2019
Cited alongside, same era.
Dense procedure captioning in narrated instructional videos
Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and Ming Zhou · 2019
Cited alongside, same era.
A unified generation-retrieval framework for image captioning
Chunpu Xu, Wei Zhao, Min Yang, Xiang Ao, Wangrong Cheng, and Jinwen Tian · 2019
Retrieval augmentation for deep neural networks
Rita Parada Ramos, Patrícia Pereira, Helena Moniz, Joao Paulo Carvalho, and Bruno Martins · 2021
Later among the works it cites.
End-to-end dense video captioning with parallel decoding
Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo · 2021
Later among the works it cites.
Open-book video captioning with retrieve-copy-generate network
Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li, Ying Deng, and Weiming Hu · 2021
Later among the works it cites.
Mugen: A playground for video-audio-text multimodal understanding and generation
Thomas Hayes, Songyang Zhang, Xi Yin, Guan Pang, Sasha Sheng, Harry Yang, Songwei Ge, Qiyuan Hu, and Devi Parikh · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Aman Chadha, Gurneet Arora, and Navpreet Kaloty · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Soda: Story oriented dense video captioning evaluation framework
Soichiro Fujita, Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, and Masaaki Nagata · 2020
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
Univl: A unified video and language pre-training model for multimodal understanding and generation, 2020
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou · 2020
Cited alongside, same era.
Swinbert: End-to-end transformers with sparse attention for video captioning
Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Later among the works it cites.
Retrieval-augmented transformer for image captioning
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid · 2022
Later among the works it cites.
Unifying event detection and captioning as sequence generation via pre-training
Qi Zhang, Yuqing Song, and Qin Jin · 2022
Later among the works it cites.
End-to-end dense video captioning as sequence generation
Wanrong Zhu, Bo Pang, Ashish V Thapliyal, William Yang Wang, and Radu Soricut · 2022
Later among the works it cites.
Retrieval augmented convolutional encoder-decoder networks for video captioning
Jingwen Chen, Yingwei Pan, Yehao Li, Ting Yao, Hongyang Chao, and Tao Mei · 2023
Later among the works it cites.
Vindlu: A recipe for effective video-and-language pretraining
Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Memory-based augmentation network for video captioning
Shuaiqi Jing, Haonan Zhang, Pengpeng Zeng, Lianli Gao, Jingkuan Song, and Heng Tao Shen · 2023
Later among the works it cites.
Retrieval-augmented image captioning
Rita Ramos, Desmond Elliott, and Bruno Martins · 2023
Later among the works it cites.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid · 2023
Later among the works it cites.