Fetching the paper…
Reading the bibliography…
In this work we present a new State-of-The-Art on the text-to-video retrieval task on MSR-VTT, LSMDC, MSVD, YouCook2 and TGIF obtained by a single model.
“VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research”, 2020
Xin Wang et al · 1904
Earlier work this paper cites.
“Making Convolutional Networks Shift-Invariant Again”
Richard Zhang · 1904
Earlier work this paper cites.
“Large-scale weakly-supervised pre-training for video action recognition”, 2019
Deepti Ghadiyaram et al · 1905
Earlier work this paper cites.
“Use What You Have: Video Retrieval Using Representations From Collaborative Experts”, 2020
Yang Liu, Samuel Albanie, Arsha Nagrani and Andrew Zisserman · 1907
Earlier work this paper cites.
“End-to-End Learning of Visual Representations from Uncurated Instructional Videos”, 2020
Antoine Miech et al · 1912
Earlier work this paper cites.
Huaishao Luo et al · 2002
Earlier work this paper cites.
“Language Models are Few-Shot Learners”, 2020
Tom. Brown et al · 2005
Earlier work this paper cites.
“AVLnet: Learning Audio-Visual Language Representations from Instructional Videos”, 2020
Andrew Rouditchenko et al · 2006
Earlier work this paper cites.
“Multi-modal Transformer for Video Retrieval”, 2020
Valentin Gabeur, Chen Sun, Karteek Alahari and Cordelia Schmid · 2007
Earlier work this paper cites.
“ImageNet: A Large-Scale Hierarchical Image Database”
J. Deng et al · 2009
Earlier work this paper cites.
“Support-set bottlenecks for video-text representation learning”, 2021
Mandela Patrick et al · 2010
Earlier work this paper cites.
“Collecting Highly Parallel Data for Paraphrase Evaluation”
David Chen and William Dolan · 2011
Earlier work this paper cites.
“COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning”
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash and Thomas Brox · 2011
Earlier work this paper cites.
“Deep Fragment Embeddings for Bidirectional Image Sentence Mapping”
Andrej Karpathy, Armand Joulin and Li Fei-Fei · 2014
Earlier work this paper cites.
“From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions”
Peter Young, Alice Lai, Micah Hodosh and Julia Hockenmaier · 2014
Earlier work this paper cites.
“Microsoft COCO Captions: Data Collection and Evaluation Server”, 2015
Xinlei Chen et al · 2015
Earlier work this paper cites.
“ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding”
Bernard Fabian Caba Victor and Juan Niebles · 2015
Earlier work this paper cites.
“TGIF: A New Dataset and Benchmark on Animated GIF Description”
Yuncheng Li et al · 2016
Cited alongside, same era.
Anna Rohrbach et al · 2016
Cited alongside, same era.
“Improved Deep Metric Learning with Multi-class N-pair Loss Objective”
Kihyuk Sohn · 2016
Cited alongside, same era.
“MSR-VTT: A Large Video Description Dataset for Bridging Video and Language”
Jun Xu, Tao Mei, Ting Yao and Yong Rui · 2016
Cited alongside, same era.
“The "something something" video database for learning and evaluating visual common sense”, 2017
Raghav Goyal et al · 2017
Cited alongside, same era.
“CNN Architectures for Large-Scale Audio Classification”, 2017
“TVQA: Localized, Compositional Video Question Answering”, 2019
Jie Lei, Licheng Yu, Mohit Bansal and Tamara. Berg · 2019
Later among the works it cites.
“W2VV++: Fully Deep Learning for Ad-hoc Video Search”, 2019
Xirong Li et al · 2019
Later among the works it cites.
“HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”
Antoine Miech et al · 2019
Later among the works it cites.
“TRECVID 2020: comprehensive campaign for evaluating video retrieval tasks across multiple application domains”
George Awad et al · 2020
Later among the works it cites.
“SEA: Sentence Encoder Assembly for Video Retrieval by Textual Queries”
Xirong Li et al · 2020
Later among the works it cites.
“Learning a Text-Video Embedding from Incomplete and Heterogeneous Data”, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shawn Hershey et al · 2017
Cited alongside, same era.
“The Kinetics Human Action Video Dataset”, 2017
Will Kay et al · 2017
Cited alongside, same era.
“End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering”, 2017
Youngjae Yu, Hyungjin Ko, Jongwook Choi and Gunhee Kim · 2017
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Cited alongside, same era.
“Predicting Visual Features From Text for Image and Video Caption Retrieval”
Jianfeng Dong, Xirong Li and Cees G.. Snoek · 2018
Cited alongside, same era.
“Learning a text-video embedding from incomplete and heterogeneous data”
Antoine Miech, Ivan Laptev and Josef Sivic · 2018
Cited alongside, same era.
“Learning joint embedding with multimodal cues for cross-modal video-text retrieval”
Niluthpol Mithun, Juncheng Li, Florian Metze and Amit Roy-Chowdhury · 2018
Cited alongside, same era.
Antoine Miech, Ivan Laptev and Josef Sivic · 2020
Later among the works it cites.
“Cross Modal Retrieval with Querybank Normalisation”, 2021
Simion-Vlad Bogolin et al · 2021
Later among the works it cites.
“Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss”, 2021
Xing Cheng et al · 2021
Later among the works it cites.
“MDMMT: Multidomain Multimodal Transformer for Video Retrieval”
Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov and Aleksandr Petiushko · 2021
Later among the works it cites.
“CLIP2Video: Mastering Video-Text Retrieval via Image CLIP”
Han Fang, Pengfei Xiong, Luhui Xu and Yu Chen · 2021
Later among the works it cites.
“CLIP2TV: An Empirical Study on Transformer-based Methods for Video-Text Retrieval”, 2021
Zijian Gao et al · 2021
Later among the works it cites.
“Lightweight Attentional Feature Fusion for Video Retrieval by Text”, 2021
Fan Hu et al · 2021
Later among the works it cites.
“Slow-Fast Auditory Streams For Audio Recognition”, 2021
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman and Dima Damen · 2021
Later among the works it cites.
“CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval”
Huaishao Luo et al · 2021
Later among the works it cites.
“A Straightforward Framework For Video Retrieval Using CLIP”, 2021
Jesúsés Portillo-Quintero, José Ortiz-Bayliss and Hugo Terashima-Marin · 2021
Later among the works it cites.
“Learning Transferable Visual Models From Natural Language Supervision”, 2021
Alec Radford et al · 2021
Later among the works it cites.
“TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment”, 2021
Jianwei Yang, Yonatan Bisk and Jianfeng Gao · 2021
Later among the works it cites.