Fetching the paper…
Reading the bibliography…
We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese.
Measuring lexical style and competence: The type-token vocabulary curve
Gilbert Youmans · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K.K. Paliwal · 1997
Earlier work this paper cites.
The part-of-speech tagging guidelines for the penn chinese treebank (3.0)
Fei Xia · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2001
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L. Chen and William B. Dolan · 2011
Earlier work this paper cites.
Hmdb: A large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso A. Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
2d human pose estimation: New benchmark and state of the art analysis
Mykhaylo Andriluka, Leonid Pishchulin, Peter V. Gehler, and Bernt Schiele · 2014
Earlier work this paper cites.
A human judgement corpus and a metric for arabic mt evaluation
Houda Bouamor, Hanan Alshikhabobakr, Behrang Mohit, and Kemal Oflazer · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
Visual semantic search: Retrieving videos via complex textual queries
Dahua Lin, Sanja Fidler, Chen Kong, and Raquel Urtasun · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Cited alongside, same era.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell · 2017
Later among the works it cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Later among the works it cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Later among the works it cites.
Fluency-guided cross-lingual image captioning
Weiyu Lan, Xirong Li, and Jianfeng Dong · 2017
Later among the works it cites.
Attention strategies for multi-source sequence-to-sequence learning
Jindřich Libovickỳ and Jindřich Helcl · 2017
Later among the works it cites.
Osu multimodal machine translation system report
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Multimodal attention for neural machine translation
Ozan Caglayan, Loïc Barrault, and Fethi Bougares · 2016
Cited alongside, same era.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia · 2016
Cited alongside, same era.
Attention-based multimodal neural machine translation
Po-Yao Huang, Frederick Liu, Sz-Rung Shiang, Jean Oh, and Chris Dyer · 2016
Cited alongside, same era.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Cited alongside, same era.
Mingbo Ma, Dapeng Li, Kai Zhao, and Liang Huang · 2017
Later among the works it cites.
Sheffield multimt: Using object posterior predictions for multimodal machine translation
Pranava Swaroop Madhyastha, Josiah Wang, and Lucia Specia · 2017
Later among the works it cites.
Reinforced video captioning with entailment rewards
Ramakanth Pasunuru and Mohit Bansal · 2017
Later among the works it cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Later among the works it cites.
Using artificial tokens to control languages for multilingual image caption generation
Satoshi Tsutsui and David Crandall · 2017
Later among the works it cites.
Video description: a survey of methods, datasets and evaluation metrics
Nayyer Aafaq, Ajmal Mian, Wei Liu, Syed Zulqarnain Gilani, and Mubarak Shah · 2018
Later among the works it cites.
nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Xinlei Chen, Rishabh Jain, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson · 2018
Later among the works it cites.
Findings of the third shared task on multimodal machine translation
Loïc Barrault, Fethi Bougares, Lucia Specia, Chiraag Lala, Desmond Elliott, and Stella Frank · 2018
Later among the works it cites.
A dataset for telling the stories of social media videos
Spandana Gella, Mike Lewis, and Marcus Rohrbach · 2018
Later among the works it cites.
The memad submission to the wmt18 multimodal translation task
Stig-Arne Grönroos, Benoit Huet, Mikko Kurimo, Jorma Laaksonen, Bernard Merialdo, Phu Pham, Mats Sjöberg, Umut Sulubacak, Jörg Tiedemann, Raphael Troncy, et al · 2018
Later among the works it cites.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Later among the works it cites.
Visual question answering dataset for bilingual image understanding: A study of cross-lingual transfer using attention maps
Nobuyuki Shimizu, Na Rong, and Takashi Miyazaki · 2018
Later among the works it cites.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang · 2018
Later among the works it cites.
Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning
Xin Wang, Yuan-Fang Wang, and William Yang Wang · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
Probing the need for visual context in multimodal machine translation
Ozan Caglayan, Pranava Madhyastha, Lucia Specia, and Loïc Barrault · 2019
Closest in time.
Coco-cn for cross-lingual image tagging, captioning and retrieval
Xirong Li, Xiaoxu Wang, Chaoxi Xu, Weiyu Lan, Qijie Wei, Gang Yang, and Jieping Xu · 2019
Closest in time.
Learning to compose topic-aware mixture of experts for zero-shot video captioning
Xin Wang, Jiawei Wu, Da Zhang, and William Yang Wang · 2019
Closest in time.
Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis · 2019
Closest in time.