Fetching the paper…
Reading the bibliography…
Employing large-scale pre-trained model CLIP to conduct video-text retrieval task (VTR) has become a new trend, which exceeds previous VTR methods.
Use What You Have: Video retrieval using representations from collaborative experts
Liu, Y.; Albanie, S.; Nagrani, A.; and Zisserman, A. ???? · 1907
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Ma, J.; Zhao, Z.; Yi, X.; Chen, J.; Hong, L.; and Chi, E. H. 2018 · 1939
Earlier work this paper cites.
A maximum entropy model for part-of-speech tagging
Ratnaparkhi, A. 1996 · 1996
Earlier work this paper cites.
Video Google: A text retrieval approach to object matching in videos
Sivic, J.; and Zisserman, A. 2003 · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Toutanova, K.; Klein, D.; Manning, C. D.; and Singer, Y. 2003 · 2003
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A.; Corrado, G.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T. 2013 · 2013
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Gong, Y.; Ke, Q.; Isard, M.; and Lazebnik, S. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Earlier work this paper cites.
The long-short story of movie description
Rohrbach, A.; Rohrbach, M.; and Schiele, B. 2015 · 2015
Earlier work this paper cites.
NII-HITACHI-UIT at TRECVID 2016
Le, D.-D.; Phan, S.; Nguyen, V.-T.; Renoust, B.; Nguyen, T. A.; Hoang, V.-N.; Ngo, T. D.; Tran, M.-T.; Watanabe, Y.; Klinkigt, M.; et al. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Cited alongside, same era.
Improving visual-semantic embeddings with hard negatives
Faghri, F.; Fleet, D.; Kiros, J.; and Fidler, S. V. 2017 · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P.; Dollár, P.; Girshick, R.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; and He, K. 2017 · 2017
Cited alongside, same era.
Query and keyframe representations for ad-hoc video search
Markatopoulou, F.; Galanopoulos, D.; Mezaris, V.; and Patras, I. 2017 · 2017
Cited alongside, same era.
Movie description
Rohrbach, A.; Torabi, A.; Rohrbach, M.; Tandon, N.; Pal, C.; Larochelle, H.; Courville, A.; and Schiele, B. 2017 · 2017
Cited alongside, same era.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020 · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Zhu, L.; and Yang, Y. 2020 · 2020
Later among the works it cites.
Vivit: A video vision transformer
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; and Schmid, C. 2021 · 2021
Closest in time.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2021 · 2021
Closest in time.
CogView: Mastering Text-to-Image Generation via Transformers
Ding, M.; Yang, Z.; Hong, W.; Zheng, W.; Zhou, C.; Yin, D.; Lin, J.; Zou, X.; Shao, Z.; Yang, H.; et al. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ruder, S. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Gu, J.; Cai, J.; Joty, S. R.; Niu, L.; and Wang, G. 2018 · 2018
Cited alongside, same era.
Squeeze-and-Excitation Networks
Hu, J.; Shen, L.; and Sun, G. 2018 · 2018
Cited alongside, same era.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Mithun, N. C.; Li, J.; Metze, F.; and Roy-Chowdhury, A. K. 2018 · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Yu, Y.; Kim, J.; and Kim, G. 2018 · 2018
Cited alongside, same era.
Closest in time.
Mdmmt: Multidomain multimodal transformer for video retrieval
Dzabraev, M.; Kalashnikov, M.; Komkov, S.; and Petiushko, A. 2021 · 2021
Closest in time.
CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
Fang, H.; Xiong, P.; Xu, L.; and Chen, Y. 2021 · 2021
Closest in time.
Hierarchical Cross-Modal Graph Consistency Learning for Video-Text Retrieval
Jin, W.; Zhao, Z.; Zhang, P.; Zhu, J.; He, X.; and Zhuang, Y. 2021 · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Lei, J.; Li, L.; Zhou, L.; Gan, Z.; Berg, T. L.; Bansal, M.; and Liu, J. 2021 · 2021
Closest in time.
Hit: Hierarchical transformer with momentum contrast for video-text retrieval
Liu, S.; Fan, H.; Qian, S.; Chen, Y.; Ding, W.; and Wang, Z. 2021 · 2021
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Closest in time.
T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
Wang, X.; Zhu, L.; and Yang, Y. 2021 · 2021
Closest in time.
VinVL: Making Visual Representations Matter in Vision-Language Models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Closest in time.