Fetching the paper…
Reading the bibliography…
In this work, we present Auto-captions on GIF, which is a new large-scale pre-training dataset for generic video understanding.
Nltk: The natural language toolkit
Edward Loper and Steven Bird · 2002
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L Chen and William B Dolan · 2011
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Earlier work this paper cites.
Click-through-based cross-view learning for image search
Yingwei Pan, Ting Yao, Tao Mei, Houqiang Li, Chong-Wah Ngo, and Yong Rui · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Sequence to sequence - video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Learning query and image similarities with ranking canonical correlation analysis
Ting Yao, Tao Mei, and Chong-Wah Ngo · 2015
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
To create what you tell: Generating videos from captions
Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei · 2017
Cited alongside, same era.
Video captioning with transferred semantic attributes
Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei · 2017
Cited alongside, same era.
Learning multimodal attention lstm networks for video captioning
Jun Xu, Ting Yao, Yongdong Zhang, and Tao Mei · 2017
Cited alongside, same era.
Incorporating copying mechanism in image captioning for learning novel objects
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2017
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Later among the works it cites.
Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning
Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian · 2019
Later among the works it cites.
Temporal deformable convolutional encoder-decoder networks for video captioning
Jingwen Chen, Yingwei Pan, Yehao Li, Ting Yao, Hongyang Chao, and Tao Mei · 2019
Later among the works it cites.
Motion guided spatial attention for video captioning
Shaoxiang Chen and Yu-Gang Jiang · 2019
Later among the works it cites.
Joint syntax representation learning and visual cue translation for video captioning
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia · 2019
Later among the works it cites.
Learning click-based deep structure-preserving embeddings with visual attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Boosting image captioning with attributes
Ting Yao, Yingwei Pan, Yehao Li, Zhaofan Qiu, and Tao Mei · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Less is more: Picking informative frames for video captioning
Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg · 2018
Cited alongside, same era.
Jointly localizing and describing events for dense video captioning
Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei · 2018
Cited alongside, same era.
Sibnet: Sibling convolutional encoder for video captioning
Sheng Liu, Zhou Ren, and Junsong Yuan · 2018
Cited alongside, same era.
Yehao Li, Yingwei Pan, Ting Yao, Hongyang Chao, Yong Rui, and Tao Mei · 2019
Later among the works it cites.
Pointing novel objects in image captioning
Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai · 2019
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Convolutional auto-encoding of sentence topics for image paragraph generation
Jing Wang, Yingwei Pan, Ting Yao, Jinhui Tang, and Tao Mei · 2019
Later among the works it cites.
Hierarchy parsing for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2019
Later among the works it cites.
X-linear attention networks for image captioning
Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei · 2020
Closest in time.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao · 2020
Closest in time.