Fetching the paper…
Reading the bibliography…
Vision-language alignment learning for video-text retrieval arouses a lot of attention in recent years.
Use what you have: Video retrieval using representations from collaborative experts
Liu, Y.; Albanie, S.; Nagrani, A.; and Zisserman, A. 2019 · 1907
Earlier work this paper cites.
The open images dataset v4
Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. 2020 · 1981
Earlier work this paper cites.
Noise estimation using density estimation for self-supervised multimodal learning
Amrani, E.; Ben-Ari, R.; Rotman, D.; and Bronstein, A. 2020 · 2003
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S.; and Zuidema, W. 2020 · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Support-set bottlenecks for video-text representation learning
Patrick, M.; Huang, P.-Y.; Asano, Y.; Metze, F.; Hauptmann, A.; Henriques, J.; and Vedaldi, A. 2020 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Robust audio-visual speech recognition under noisy audio-video conditions
Stewart, D.; Seymour, R.; Pass, A.; and Ming, J. 2013 · 2013
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R.; Salakhutdinov, R.; and Zemel, R. S. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Venugopalan, S.; Xu, H.; Donahue, J.; Rohrbach, M.; Mooney, R.; and Saenko, K. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
The long-short story of movie description
Rohrbach, A.; Rohrbach, M.; and Schiele, B. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Cited alongside, same era.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Xie, S.; Sun, C.; Huang, J.; Tu, Z.; and Murphy, K. 2018 · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Yu, Y.; Kim, J.; and Kim, G. 2018 · 2018
Cited alongside, same era.
Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video Retrieval
Chen, Q.; Liu, Y.; and Albanie, S. 2021 · 2021
Later among the works it cites.
Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss
Cheng, X.; Lin, H.; Wu, X.; Yang, F.; and Shen, D. 2021 · 2021
Later among the works it cites.
Teachtext: Crossmodal generalized distillation for text-video retrieval
Croitoru, I.; Bogolin, S.-V.; Leordeanu, M.; Jin, H.; Zisserman, A.; Albanie, S.; and Liu, Y. 2021 · 2021
Later among the works it cites.
Mdmmt: Multidomain multimodal transformer for video retrieval
Dzabraev, M.; Kalashnikov, M.; Komkov, S.; and Petiushko, A. 2021 · 2021
Later among the works it cites.
Clip2video: Mastering video-text retrieval via image clip
Fang, H.; Xiong, P.; Xu, L.; and Chen, Y. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cross-modal and hierarchical modeling of video and text
Zhang, B.; Hu, H.; and Sha, F. 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
Deep supervised cross-modal retrieval
Zhen, L.; Hu, P.; Wang, X.; and Peng, D. 2019 · 2019
Cited alongside, same era.
Multi-modal transformer for video retrieval
Gabeur, V.; Sun, C.; Alahari, K.; and Schmid, C. 2020 · 2020
Cited alongside, same era.
KeyBERT: Minimal keyword extraction with BERT
Grootendorst, M. 2020 · 2020
Cited alongside, same era.
Actbert: Learning global-local video-text representations
Zhu, L.; and Yang, Y. 2020 · 2020
Cited alongside, same era.
Multi-Feature Graph Attention Network for Cross-Modal Video-Text Retrieval
Hao, X.; Zhou, Y.; Wu, D.; Zhang, W.; Li, B.; and Wang, W. 2021 · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Lei, J.; Li, L.; Zhou, L.; Gan, Z.; Berg, T. L.; Bansal, M.; and Liu, J. 2021 · 2021
Later among the works it cites.
Hit: Hierarchical transformer with momentum contrast for video-text retrieval
Liu, S.; Fan, H.; Qian, S.; Chen, Y.; Ding, W.; and Wang, Z. 2021 · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2021 · 2021
Later among the works it cites.
A straightforward framework for video retrieval using clip
Portillo-Quintero, J. A.; Ortiz-Bayliss, J. C.; and Terashima-Marín, H. 2021 · 2021
Later among the works it cites.
Dual adversarial graph neural networks for multi-label cross-modal retrieval
Qian, S.; Xue, D.; Zhang, H.; Fang, Q.; and Xu, C. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
T2vlad: global-local sequence alignment for text-video retrieval
Wang, X.; Zhu, L.; and Yang, Y. 2021 · 2021
Later among the works it cites.
Taco: Token-aware cascade contrastive learning for video-text alignment
Yang, J.; Bisk, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
Visual Consensus Modeling for Video-Text Retrieval
Cao, S.; Wang, B.; Zhang, W.; and Ma, L. 2022 · 2022
Later among the works it cites.