Fetching the paper…
Reading the bibliography…
Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries.
Collecting highly parallel data for paraphrase evaluation
David Chen and William B Dolan. 2011 · 2011
Earlier work this paper cites.
Slow feature analysis for human action recognition
Zhang Zhang and Dacheng Tao. 2012 · 2012
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. 2020 · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020 · 2020
Earlier work this paper cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020 · 2020
Earlier work this paper cites.
On pursuit of designing multi-modal transformer for video grounding
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021a · 2021
Earlier work this paper cites.
Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss
Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen. 2021 · 2021
Earlier work this paper cites.
Teachtext: Crossmodal generalized distillation for text-video retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021 · 2021
Earlier work this paper cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
Training neural networks with fixed sparse masks
Yi-Lin Sung, Varun Nair, and Colin A Raffel. 2021 · 2021
Cited alongside, same era.
T2vlad: global-local sequence alignment for text-video retrieval
Xiaohan Wang, Linchao Zhu, and Yi Yang. 2021 · 2021
Cited alongside, same era.
Taco: Token-aware cascade contrastive learning for video-text alignment
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. 2021 · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021 · 2021
Cited alongside, same era.
Learning visual representation from modality-shared contrastive language-image pre-training
Haoxuan You, Luowei Zhou, Bin Xiao, Noel Codella, Yu Cheng, Ruochen Xu, Shih-Fu Chang, and Lu Yuan. 2022 · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022 · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022 · 2022
Later among the works it cites.
Unsupervised pre-training for temporal action localization tasks
Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao, Jue Wang, and Yuexian Zou. 2022 · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. 2022 · 2022
Cited alongside, same era.
X-pool: Cross-modal language-video attention for text-video retrieval
Satya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. 2022 · 2022
Cited alongside, same era.
Visual relation-aware unsupervised video captioning
Puzhao Ji, Meng Cao, and Yuexian Zou. 2022 · 2022
Cited alongside, same era.
Visual prompt tuning
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022 · 2022
Cited alongside, same era.
Convolutional bypasses are better vision transformer adapters
Shibo Jie and Zhi-Hong Deng. 2022 · 2022
Cited alongside, same era.
Expectation-maximization contrastive learning for compact video-and-language representations
Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David Clifton, and Jie Chen. 2022 · 2022
Cited alongside, same era.
Scaling & shifting your features: A new baseline for efficient model tuning
Dongze Lian, Daquan Zhou, Jiashi Feng, and Xinchao Wang. 2022 · 2022
Cited alongside, same era.
Iterative proposal refinement for weakly-supervised video grounding
Meng Cao, Fangyun Wei, Can Xu, Xiubo Geng, Long Chen, Can Zhang, Yuexian Zou, Tao Shen, and Daxin Jiang. 2023 · 2023
Later among the works it cites.
Rgnet: A unified retrieval and grounding network for long videos
Tanveer Hannan, Md Mohaiminul Islam, Thomas Seidl, and Gedas Bertasius. 2023 · 2023
Later among the works it cites.
Vop: Text-video co-operative prompt tuning for cross-modal retrieval
Siteng Huang, Biao Gong, Yulin Pan, Jianwen Jiang, Yiliang Lv, Yuyuan Li, and Donglin Wang. 2023 · 2023
Later among the works it cites.
Video referring expression comprehension via transformer with content-conditioned query
Jiang Ji, Meng Cao, Tengtao Song, Long Chen, Yi Wang, and Yuexian Zou. 2023 · 2023
Later among the works it cites.
Zeroi2v: Zero-cost adaptation of pre-trained transformers from image to video
Xinhao Li and Limin Wang. 2023 · 2023
Later among the works it cites.
Uniadapter: Unified parameter-efficient transfer learning for cross-modal modeling
Haoyu Lu, Mingyu Ding, Yuqi Huo, Guoxing Yang, Zhiwu Lu, Masayoshi Tomizuka, and Wei Zhan. 2023 · 2023
Later among the works it cites.
Improving reference-based distinctive image captioning with contrastive rewards
Yangjun Mao, Jun Xiao, Dong Zhang, Meng Cao, Jian Shao, Yueting Zhuang, and Long Chen. 2023 · 2023
Later among the works it cites.
Video-text retrieval by supervised multi-space multi-grained alignment
Yimu Wang and Peng Shi. 2023 · 2023
Later among the works it cites.
Unified coarse-to-fine alignment for video-text retrieval
Ziyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Concept-aware video captioning: Describing videos with effective prior information
Bang Yang, Meng Cao, and Yuexian Zou. 2023 · 2023
Later among the works it cites.
Side4video: Spatial-temporal side network for memory-efficient image-to-video transfer learning
Huanjin Yao, Wenhao Wu, and Zhiheng Li. 2023 · 2023
Later among the works it cites.
Qilin-med: Multi-stage knowledge injection advanced medical large language model
Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou, Yining Hua, and Andrew Liu. 2023 · 2023
Later among the works it cites.
Multimodal video adapter for parameter efficient video text retrieval
Bowen Zhang, Xiaojie Jin, Weibo Gong, Kai Xu, Zhao Zhang, Peng Wang, Xiaohui Shen, and Jiashi Feng. 2023 · 2023
Later among the works it cites.
Unipt: Universal parallel tuning for transfer learning with efficient parameter and memory
Haiwen Diao, Bo Wan, Ying Zhang, Xu Jia, Huchuan Lu, and Long Chen. 2024 · 2024
Closest in time.
Exploiting auxiliary caption for video grounding
Hongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li, Zhihong Zhu, and Yuexian Zou. 2024 · 2024
Closest in time.