Fetching the paper…
Reading the bibliography…
We present a new state-of-the-art on the text to video retrieval task on MSRVTT and LSMDC benchmarks where our model outperforms all previous solutions by a large margin.
“VideoBERT: A Joint Model for Video and Language Representation Learning”, 2019
Chen Sun et al · 1904
Earlier work this paper cites.
“Video Classification with Channel-Separated Convolutional Networks”, 2019
Du Tran, Heng Wang, Lorenzo Torresani and Matt Feiszli · 1904
Earlier work this paper cites.
“Large-scale weakly-supervised pre-training for video action recognition”, 2019
Deepti Ghadiyaram et al · 1905
Earlier work this paper cites.
“Use What You Have: Video Retrieval Using Representations From Collaborative Experts”, 2020
Yang Liu, Samuel Albanie, Arsha Nagrani and Andrew Zisserman · 1907
Earlier work this paper cites.
“Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings”, 2019
Michael Wray, Diane Larlus, Gabriela Csurka and Dima Damen · 1908
Earlier work this paper cites.
“End-to-End Learning of Visual Representations from Uncurated Instructional Videos”, 2020
Antoine Miech et al · 1912
Earlier work this paper cites.
“AVLnet: Learning Audio-Visual Language Representations from Instructional Videos”, 2020
Andrew Rouditchenko et al · 2006
Earlier work this paper cites.
“Multi-modal Transformer for Video Retrieval”, 2020
Valentin Gabeur, Chen Sun, Karteek Alahari and Cordelia Schmid · 2007
Earlier work this paper cites.
“ImageNet: A Large-Scale Hierarchical Image Database”
J. Deng et al · 2009
Earlier work this paper cites.
“Support-set bottlenecks for video-text representation learning”, 2021
Mandela Patrick et al · 2010
Earlier work this paper cites.
“A Short Note on the Kinetics-700-2020 Human Action Dataset”, 2020
Lucas Smaira et al · 2010
Earlier work this paper cites.
“Collecting Highly Parallel Data for Paraphrase Evaluation”
David Chen and William Dolan · 2011
Earlier work this paper cites.
“Deep Fragment Embeddings for Bidirectional Image Sentence Mapping”
Andrej Karpathy, Armand Joulin and Li Fei-Fei · 2014
Earlier work this paper cites.
“Large-scale Video Classification with Convolutional Neural Networks”
Andrej Karpathy et al · 2014
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Cited alongside, same era.
“A Dataset for Movie Description”, 2015
Anna Rohrbach, Marcus Rohrbach, Niket Tandon and Bernt Schiele · 2015
Cited alongside, same era.
“Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research”, 2015
Atousa Torabi, Christopher Pal, Hugo Larochelle and Aaron Courville · 2015
Cited alongside, same era.
“TGIF: A New Dataset and Benchmark on Animated GIF Description”
Yuncheng Li et al · 2016
Cited alongside, same era.
Anna Rohrbach et al · 2016
“Predicting Visual Features From Text for Image and Video Caption Retrieval”
Jianfeng Dong, Xirong Li and Cees G.. Snoek · 2018
Later among the works it cites.
“Exploring the Limits of Weakly Supervised Pretraining”
Dhruv Mahajan et al · 2018
Later among the works it cites.
“Learning joint embedding with multimodal cues for cross-modal video-text retrieval”
Niluthpol Mithun, Juncheng Li, Florian Metze and Amit Roy-Chowdhury · 2018
Later among the works it cites.
“A Closer Look at Spatiotemporal Convolutions for Action Recognition”, 2018
Du Tran et al · 2018
Later among the works it cites.
Saining Xie et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Learning Language-Visual Embedding for Movie Understanding with Natural-Language”, 2016
Atousa Torabi, Niket Tandon and Leonid Sigal · 2016
Cited alongside, same era.
“MSR-VTT: A Large Video Description Dataset for Bridging Video and Language”
Jun Xu, Tao Mei, Ting Yao and Yong Rui · 2016
Cited alongside, same era.
“The "something something" video database for learning and evaluating visual common sense”, 2017
Raghav Goyal et al · 2017
Cited alongside, same era.
“CNN Architectures for Large-Scale Audio Classification”, 2017
Shawn Hershey et al · 2017
Cited alongside, same era.
“The Kinetics Human Action Video Dataset”, 2017
Will Kay et al · 2017
Cited alongside, same era.
“Near-duplicate video retrieval with deep metric learning”
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras and Yiannis Kompatsiaris · 2017
Cited alongside, same era.
“Dense-Captioning Events in Videos”
Ranjay Krishna et al · 2017
Cited alongside, same era.
Youngjae Yu, Jongseok Kim and Gunhee Kim · 2018
Later among the works it cites.
“Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction”, 2018
Luowei Zhou, Nathan Louis and Jason. Corso · 2018
Later among the works it cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Later among the works it cites.
“Dual Encoding for Zero-Example Video Retrieval”, 2019
Jianfeng Dong et al · 2019
Later among the works it cites.
“SlowFast Networks for Video Recognition”, 2019
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik and Kaiming He · 2019
Later among the works it cites.
“HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”
Antoine Miech et al · 2019
Later among the works it cites.
“TRECVID 2020: comprehensive campaign for evaluating video retrieval tasks across multiple application domains”
George Awad et al · 2020
Later among the works it cites.
“Learning a Text-Video Embedding from Incomplete and Heterogeneous Data”, 2020
Antoine Miech, Ivan Laptev and Josef Sivic · 2020
Later among the works it cites.
“A Straightforward Framework For Video Retrieval Using CLIP”, 2021
Jesúsés Portillo-Quintero, José Ortiz-Bayliss and Hugo Terashima-Marín · 2021
Closest in time.