Fetching the paper…
Reading the bibliography…
We address the problem of cross-modal fine-grained action retrieval between text and video.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2013
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Learning fine-grained image similarity with deep ranking
Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jingbin Wang, James Philbin, Bo Chen, and Ying Wu · 2014
Earlier work this paper cites.
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon · 2015
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J Corso · 2015
Earlier work this paper cites.
Zero-shot learning via semantic similarity embedding
Ziming Zhang and Venkatesh Saligrama · 2015
Earlier work this paper cites.
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
An empirical study and analysis of generalized zero-shot learning for object recognition in the wild
Wei-Lun Chao, Soravit Changpinyo, Boqing Gong, and Fei Sha · 2016
Earlier work this paper cites.
Deep image retrieval: Learning global representations for image search
Albert Gordo, Jon Almazán, Jérome Revaud, and Diane Larlus · 2016
Earlier work this paper cites.
Learning joint representations of videos and sentences with web image search
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, and Naokazu Yokoya · 2016
Earlier work this paper cites.
CNN image retrieval learns from BoW: Unsupervised fine-tuning with hard examples
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2016
Earlier work this paper cites.
Recognizing fine-grained and composite activities using hand-centric features and script data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, and Bernt Schiele · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Cited alongside, same era.
Learning language-visual embedding for movie understanding with natural-language
Atousa Torabi, Niket Tandon, and Leonid Sigal · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Predicting visual features from text for image and video caption retrieval
Jianfeng Dong, Xirong Li, and Cees GM Snoek · 2018
Later among the works it cites.
Dual dense encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, and Xun Wang · 2018
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Later among the works it cites.
Action2vec: A crossmodal embedding approach to action learning
Meera Hahn, Andrew Silva, and James M. Rehg · 2018
Later among the works it cites.
On the effectiveness of task granularity for transfer learning
Farzaneh Mahdisoltani, Guillaume Berger, Waseem Gharbieh, David Fleet, and Roland Memisevic · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Beyond triplet loss: a deep quadruplet network for person re-identification
Weihua Chen, Xiaotang Chen, Ianguo Zhang, and Kaiqi Huang · 2017
Cited alongside, same era.
Beyond instance-level image retrieval: Leveraging captions to learn a global visual representation for semantic retrieval
Albert Gordo and Diane Larlus · 2017
Cited alongside, same era.
The “Something Something” Video Database for Learning and Evaluating Visual Common Sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Cited alongside, same era.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe · 2017
Cited alongside, same era.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Cited alongside, same era.
Enhancing video summarization via vision-language embedding
Bryan A Plummer, Matthew Brown, and Svetlana Lazebnik · 2017
Cited alongside, same era.
Later among the works it cites.
Learning a text-video embedding from incomplete and heterogeneous data
Antoine Miech, Ivan Laptev, and Josef Sivic · 2018
Later among the works it cites.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K Roy-Chowdhury · 2018
Later among the works it cites.
High order neural networks for video classification
Jie Shao, Kai Hu, Yixin Bao, Yining Lin, and Xiangyang Xue · 2018
Later among the works it cites.
Cross modal embeddings for video and audio retrieval
Didac Surís, Amanda Duarte, Amaia Salvador, Jordi Torres, and Xavier Giróo-i Nieto · 2018
Later among the works it cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Wang Liwei, Li Yin, Huang Jing, and Svetlana Lazebnik · 2019
Closest in time.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Closest in time.
Learning visual actions using multiple verb-only labels
Michael Wray and Dima Damen · 2019
Closest in time.