Fetching the paper…
Reading the bibliography…
Text-Video Retrieval (TVR) aims to align and associate relevant video content with corresponding natural language queries.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Yu, Y.; Kim, J.; and Kim, G. 2018 · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Earlier work this paper cites.
Multi-modal transformer for video retrieval
Gabeur, V.; Sun, C.; Alahari, K.; and Schmid, C. 2020 · 2020
Earlier work this paper cites.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Lei, J.; Yu, L.; Berg, T. L.; and Bansal, M. 2020 · 2020
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Earlier work this paper cites.
On pursuit of designing multi-modal transformer for video grounding
Cao, M.; Chen, L.; Shou, M. Z.; Zhang, C.; and Zou, Y. 2021 · 2021
Earlier work this paper cites.
Improving video retrieval by adaptive margin
He, F.; Wang, Q.; Feng, Z.; Jiang, W.; Lü, Y.; Zhu, Y.; and Tan, X. 2021 · 2021
Earlier work this paper cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Lei, J.; Li, L.; Zhou, L.; Gan, Z.; Berg, T. L.; Bansal, M.; and Liu, J. 2021 · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
RR-Net: Relation reasoning for end-to-end human-object interaction detection
Yang, D.; Zou, Y.; Zhang, C.; Cao, M.; and Chen, J. 2021 · 2021
Earlier work this paper cites.
Cola: Weakly-supervised temporal action localization with snippet contrastive learning
Zhang, C.; Cao, M.; Yang, D.; Chen, J.; and Zou, Y. 2021 · 2021
Cited alongside, same era.
Vision transformer adapter for dense predictions
Chen, Z.; Duan, Y.; Wang, W.; He, J.; Lu, T.; Dai, J.; and Qiao, Y. 2022 · 2022
Cited alongside, same era.
Ms-tct: Multi-scale temporal convtransformer for action detection
Dai, R.; Das, S.; Kahatapitiya, K.; Ryoo, M. S.; and Brémond, F. 2022 · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Cited alongside, same era.
X-pool: Cross-modal language-video attention for text-video retrieval
Gorti, S. K.; Vouitsis, N.; Ma, J.; Golestan, K.; Volkovs, M.; Garg, A.; and Yu, G. 2022 · 2022
Cited alongside, same era.
Video Referring Expression Comprehension via Transformer with Content-conditioned Query
Ji, J.; Cao, M.; Song, T.; Chen, L.; Wang, Y.; and Zou, Y. 2023 · 2023
Later among the works it cites.
Unified coarse-to-fine alignment for video-text retrieval
Wang, Z.; Sung, Y.-L.; Cheng, F.; Bertasius, G.; and Bansal, M. 2023 · 2023
Later among the works it cites.
Concept-aware video captioning: Describing videos with effective prior information
Yang, B.; Cao, M.; and Zou, Y. 2023 · 2023
Later among the works it cites.
Qilin-med: Multi-stage knowledge injection advanced medical large language model
Ye, Q.; Liu, J.; Chong, D.; Zhou, P.; Hua, Y.; and Liu, A. 2023 · 2023
Later among the works it cites.
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
Cao, M.; Tang, H.; Huang, J.; Jin, P.; Zhang, C.; Liu, R.; Chen, L.; Liang, X.; Yuan, L.; and Li, G. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tencent text-video retrieval: Hierarchical cross-modal interactions with multi-level representations
Jiang, J.; Min, S.; Kong, W.; Wang, H.; Li, Z.; and Liu, W. 2022 · 2022
Cited alongside, same era.
Expectation-maximization contrastive learning for compact video-and-language representations
Jin, P.; Huang, J.; Liu, F.; Wu, X.; Ge, S.; Song, G.; Clifton, D.; and Chen, J. 2022 · 2022
Cited alongside, same era.
Exploring plain vision transformer backbones for object detection
Li, Y.; Mao, H.; Girshick, R.; and He, K. 2022 · 2022
Cited alongside, same era.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2022 · 2022
Cited alongside, same era.
X-clip: End-to-end multi-grained contrastive learning for video-text retrieval
Ma, Y.; Xu, G.; Sun, X.; Yan, M.; Zhang, J.; and Ji, R. 2022 · 2022
Cited alongside, same era.
I 2 Transformer: Intra-and inter-relation embedding transformer for TV show captioning
Tu, Y.; Li, L.; Su, L.; Gao, S.; Yan, C.; Zha, Z.-J.; Yu, Z.; and Huang, Q. 2022 · 2022
Cited alongside, same era.
Disentangled representation learning for text-video retrieval
Wang, Q.; Zhang, Y.; Zheng, Y.; Pan, P.; and Hua, X.-S. 2022 · 2022
Cited alongside, same era.
Closest in time.
Localmamba: Visual state space model with windowed selective scan
Huang, T.; Pei, X.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2024 · 2024
Closest in time.
Textual Inversion and Self-supervised Refinement for Radiology Report Generation
Luo, Y.; Li, H.; Wu, X.; Cao, M.; Huang, X.; Zhu, Z.; Liao, P.; Chen, H.; and Zhang, Y. 2024 · 2024
Closest in time.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
Ma, J.; Li, F.; and Wang, B. 2024 · 2024
Closest in time.
Efficientvmamba: Atrous selective scan for light weight visual mamba
Pei, X.; Huang, T.; and Xu, C. 2024 · 2024
Closest in time.
Gamba: Marry gaussian splatting with mamba for single view 3d reconstruction
Shen, Q.; Yi, X.; Wu, Z.; Zhou, P.; Zhang, H.; Yan, S.; and Wang, X. 2024 · 2024
Closest in time.
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
Wang, J.; Sun, G.; Wang, P.; Liu, D.; Dianat, S.; Rabbani, M.; Rao, R.; and Tao, Z. 2024 · 2024
Closest in time.
Uncertainty-aware sign language video retrieval with probability distribution modeling
Wu, X.; Li, H.; Luo, Y.; Cheng, X.; Zhuang, X.; Cao, M.; and Fu, K. 2024 · 2024
Closest in time.
Plainmamba: Improving non-hierarchical mamba in visual recognition
Yang, C.; Chen, Z.; Espinosa, M.; Ericsson, L.; Wang, Z.; Liu, J.; and Crowley, E. J. 2024 · 2024
Closest in time.
MambaOut: Do We Really Need Mamba for Vision?
Yu, W.; and Wang, X. 2024 · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; and Wang, X. 2024 · 2024
Closest in time.