Fetching the paper…
Reading the bibliography…
In this work, we tackle the problem of text-to-video retrieval (T2VR).
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher · 2009
Earlier work this paper cites.
Collecting Highly Parallel Data for Paraphrase Evaluation
David Chen and William Dolan · 2011
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Localizing Moments in Video with Natural Language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Multi-task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
Roberto Cipolla, Yarin Gal, and Alex Kendall · 2018
Earlier work this paper cites.
Using the borda rule for ranking sets of objects
Andreas Darmann and Christian Klamler · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
VideoBERT: A Joint Model for Video and Language Representation Learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang · 2019
Earlier work this paper cites.
Debiased Contrastive Learning
Ching-Yao Chuang, Joshua Robinson, Lin Yen-Chen, Antonio Torralba, and Stefanie Jegelka · 2020
Earlier work this paper cites.
Multi-modal Transformer for Video Retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Earlier work this paper cites.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab and Matei Zaharia · 2020
Earlier work this paper cites.
Use What You Have: Video Retrieval Using Representations From Collaborative Experts, 2020
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2020
Earlier work this paper cites.
Document Ranking with a Pretrained Sequence-to-Sequence Model, 2020
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin · 2020
Earlier work this paper cites.
A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation, 2020
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen · 2020
Cited alongside, same era.
ActBERT: Learning Global-Local Video-Text Representations
Linchao Zhu and Yi Yang · 2020
Cited alongside, same era.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss, 2021
Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Later among the works it cites.
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji · 2022
Later among the works it cites.
Expanding Language-Image Pretrained Models for General Video Recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Later among the works it cites.
Long-form video-language pre-training with multimodal temporal contrastive learning
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu · 2022
Later among the works it cites.
Disentangled Representation Learning for Text-Video Retrieval, 2022
Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan, and Xian-Sheng Hua · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A white box analysis of colbert
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant · 2021
Cited alongside, same era.
Less is More: CLIPBERT for Video-and-Language Learning via Sparse Sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, Joao Henriques, and Andrea Vedaldi · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
Xiaohan Wang, Linchao Zhu, and Yi Yang · 2021
Cited alongside, same era.
On Semantic Similarity in Video Retrieval
Michael Wray, Hazel Doughty, and Dima Damen · 2021
Cited alongside, same era.
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Cited alongside, same era.
FILIP: Fine-grained Interactive Language-Image Pre-Training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Later among the works it cites.
LiT: Zero-Shot Transfer with Locked-image text Tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Later among the works it cites.
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang · 2022
Later among the works it cites.
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, and Jing Liu · 2023
Later among the works it cites.
Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval
Chaorui Deng, Qi Chen, Pengda Qin, Da Chen, and Qi Wu · 2023
Later among the works it cites.
PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval
Peiyan Guan, Renjing Pei, Bin Shao, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu, Songcen Xu, Youliang Yan, and Edmund Y. Lam · 2023
Later among the works it cites.
Multiple Tasks for Multiple Objectives: A New Multiobjective Optimization Method via Multitask Optimization
Jian-Yu Li, Zhi-Hui Zhan, Yun Li, and Jun Zhang · 2023
Later among the works it cites.
Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?
Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang · 2023
Later among the works it cites.
CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment
Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
Meng Cao, Haoran Tang, Jinfa Huang, Peng Jin, Can Zhang, Ruyang Liu, Long Chen, Xiaodan Liang, Li Yuan, and Ge Li · 2024
Later among the works it cites.
The MSR-Video to Text Dataset with Clean Annotations
Haoran Chen, Jianmin Li, Simone Frintrop, and Xiaolin Hu · 2024
Later among the works it cites.
ColPali: Efficient Document Retrieval with Vision Language Models, 2024
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo · 2024
Later among the works it cites.
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao · 2024
Later among the works it cites.
VidLA: Video-Language Alignment at Scale
Mamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran, Benjamin Z. Yao, Belinda Zeng, Mubarak Shah, and Trishul Chilimbi · 2024
Later among the works it cites.