Dual dense encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, and Xun Wang · 2019
Later among the works it cites.
A comprehensive empirical comparison of hubness reduction in high-dimensional spaces
Roman Feldbauer and Arthur Flexer · 2019
Later among the works it cites.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Later among the works it cites.
Squeeze-and-excitation networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2019
Later among the works it cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Later among the works it cites.
Abstractive summarization of reddit posts with multi-level memory networks
Original
Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim · 2019
Later among the works it cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Later among the works it cites.
A strong and robust baseline for text-image matching
Original
Fangyu Liu and Rongtian Ye · 2019
Later among the works it cites.
Use what you have: Video retrieval using representations from collaborative experts
Original
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Multi-similarity loss with general pair weighting for deep metric learning
Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R Scott · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Eric Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan fang Wang, and William Yang Wang · 2019
Later among the works it cites.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen · 2019
Later among the works it cites.
Coloring with limited data: Few-shot colorization via memory augmented networks
Seungjoo Yoo, Hyojin Bahng, Sunghyo Chung, Junsoo Lee, Jaehyuk Chang, and Jaegul Choo · 2019
Later among the works it cites.
Smooth-ap: Smoothing the path towards large-scale image retrieval
Andrew Brown, Weidi Xie, Vicky Kalogeiton, and Andrew Zisserman · 2020
Later among the works it cites.
Team ruc ai.m3: Technical report in video pentathlon challenge 2020
Shizhe Chen, Yida Zhao, and Qin Jin · 2020
Later among the works it cites.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Later among the works it cites.
Meshed-memory transformer for image captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara · 2020
Later among the works it cites.
scikit-hubness: Hubness reduction and approximate neighbor search
Roman Feldbauer, Thomas Rattei, and Arthur Flexer · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Xiaowei Hu, Pengchuan Zhang, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Later among the works it cites.
Hal: Improved text-image matching by mitigating visual semantic hubs
Original
Fangyu Liu, Rongtian Ye, Xun Wang, and Shuaipeng Li · 2020
Later among the works it cites.
Queryd: a video dataset with high-quality textual and audio narrations
Original
Andreea-Maria Oncescu, Joao F. Henriques, Yang Liu, Andrew Zisserman Zisserman, and Samuel Albanie · 2020
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
Original
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, João Henriques, and Andrea Vedaldi · 2020
Later among the works it cites.
Revisiting training strategies and generalization performance in deep metric learning
Karsten Roth, Timo Milbich, Samarth Sinha, Prateek Gupta, Bjorn Ommer, and Joseph Paul Cohen · 2020
Later among the works it cites.
Cross-batch memory for embedding learning
Xun Wang, Haozhi Zhang, Weilin Huang, and Matthew R Scott · 2020
Later among the works it cites.
Interactive key-value memory-augmented attention for image paragraph captioning
Chunpu Xu, Yu Li, Chengming Li, Xiang Ao, Min Yang, and Jinwen Tian · 2020
Later among the works it cites.
Ruc_aim3 at trecvid 2020: Ad-hoc video search & video to text description
Yida Zhao, Yuqing Song, Shizhe Chen, and Qin Jin · 2020
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Closest in time.
Improving memory banks for unsupervised learning with large mini-batch, consistency and hard negative mining
Adrian Bulat, Enrique Sánchez-Lozano, and Georgios Tzimiropoulos · 2021
Closest in time.
Teachtext: Crossmodal generalized distillation for text-video retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu · 2021
Closest in time.
Clip2video: Mastering video-text retrieval via image clip
Original
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Closest in time.
Retrieve fast, rerank smart: Cooperative and joint approaches for improved cross-modal retrieval
Original
Gregor Geigle, Jonas Pfeiffer, Nils Reimers, Ivan Vulić, and Iryna Gurevych · 2021
Closest in time.
Rethinking preventing class-collapsing in metric learning with margin-based losses
Elad Levi, Tete Xiao, Xiaolong Wang, and Trevor Darrell · 2021
Closest in time.
Adaptive cross-modal prototypes for cross-domain visual-language retrieval
Yang Liu, Qingchao Chen, and Samuel Albanie · 2021
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Original
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2021
Closest in time.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2021
Closest in time.
Domain adaptation in multi-view embedding for cross-modal video retrieval
Original
Jonathan Munro, Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen · 2021
Closest in time.
Audio retrieval with natural language queries
Andreea-Maria Oncescu, A Koepke, João F Henriques, Zeynep Akata, and Samuel Albanie · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Original
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Vinvl: Making visual representations matter in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Closest in time.