Fetching the paper…
Reading the bibliography…
Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions.
Lifting the curse of dimensionality
Frances Y Kuo and Ian H Sloan · 2005
Earlier work this paper cites.
Learning deep architectures for ai
Yoshua Bengio et al · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William B Dolan · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Disentangled representation learning gan for pose-invariant face recognition
Luan Tran, Xi Yin, and Xiaoming Liu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A framework for the quantitative evaluation of disentangled representations
Cian Eastwood and Christopher KI Williams · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning deep disentangled embeddings with the f-statistic loss
Karl Ridgeway and Michael C Mozer · 2018
Earlier work this paper cites.
Theory and evaluation metrics for learning disentangled representations
Kien Do and Truyen Tran · 2019
Earlier work this paper cites.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2019
Earlier work this paper cites.
Disentangling factors of variation using few labels
Francesco Locatello, Michael Tschannen, Stefan Bauer, Gunnar Rätsch, Bernhard Schölkopf, and Olivier Bachem · 2019
Earlier work this paper cites.
Disentangled graph convolutional networks
Jianxin Ma, Peng Cui, Kun Kuang, Xin Wang, and Wenwu Zhu · 2019
Earlier work this paper cites.
Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness
Raphael Suter, Djordje Miladinovic, Bernhard Schölkopf, and Stefan Bauer · 2019
Cited alongside, same era.
Are disentangled representations helpful for abstract visual reasoning?
Sjoerd Van Steenkiste, Francesco Locatello, Jürgen Schmidhuber, and Olivier Bachem · 2019
Cited alongside, same era.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Gabriela Csurka, Diane Larlus, and Dima Damen · 2019
Cited alongside, same era.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Cited alongside, same era.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Cross modal retrieval with querybank normalisation
Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu, and Samuel Albanie · 2022
Later among the works it cites.
X-pool: Cross-modal language-video attention for text-video retrieval
Satya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu · 2022
Later among the works it cites.
Expectation-maximization contrastive learning for compact video-and-language representations
Peng Jin, Jinfa Huang, Fenglin Liu, Xian Wu, Shen Ge, Guoli Song, David Clifton, and Jie Chen · 2022
Later among the works it cites.
Toward 3d spatial reasoning for human-like text-based visual question answering
Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen · 2022
Later among the works it cites.
Joint learning of object graph and relation graph for visual question answering
Hao Li, Xu Li, Belhal Karimi, Jie Chen, and Mingming Sun · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Teachtext: Crossmodal generalized distillation for text-video retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Cited alongside, same era.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Adaptive cross-modal prototypes for cross-domain visual-language retrieval
Yang Liu, Qingchao Chen, and Samuel Albanie · 2021
Cited alongside, same era.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G Hauptmann, Joao F. Henriques, and Andrea Vedaldi · 2021
Cited alongside, same era.
Locality guidance for improving vision transformers on tiny datasets
Kehan Li, Runyi Yu, Zhennan Wang, Li Yuan, Guoli Song, and Jie Chen · 2022
Later among the works it cites.
Ts2-net: Token shift and selection transformer for text-video retrieval
Yuqi Liu, Pengfei Xiong, Luhui Xu, Shengming Cao, and Qin Jin · 2022
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Later among the works it cites.
Disentangled representation learning for text-video retrieval
Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan, and Xian-Sheng Hua · 2022
Later among the works it cites.
Clip-vip: Adapting pre-trained image-text model to video-language representation alignment
Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo · 2022
Later among the works it cites.
M-mix: Generating hard negatives via multi-sample mixing for contrastive learning
Shaofeng Zhang, Meng Liu, Junchi Yan, Hengrui Zhang, Lingxiao Huang, Xiaokang Yang, and Pinyan Lu · 2022
Later among the works it cites.
Align representations with base: A new approach to self-supervised learning
Shaofeng Zhang, Lyn Qiu, Feng Zhu, Junchi Yan, Hengrui Zhang, Rui Zhao, Hongyang Li, and Xiaokang Yang · 2022
Later among the works it cites.
Parallel vertex diffusion for unified visual grounding
Zesen Cheng, Kehan Li, Peng Jin, Xiangyang Ji, Li Yuan, Chang Liu, and Jie Chen · 2023
Closest in time.
Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning
Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, and Jie Chen · 2023
Closest in time.
Diffusionret: Generative text-video retrieval with diffusion model
Peng Jin, Hao Li, Zesen Cheng, Kehan Li, Xiangyang Ji, Chang Liu, Li Yuan, and Jie Chen · 2023
Closest in time.
Fits: Fine-grained two-stage training for knowledge-aware question answering
Qichen Ye, Bowen Cao, Nuo Chen, Weiyuan Xu, and Yuexian Zou · 2023
Closest in time.
Shaofeng Zhang, Feng Zhu, Rui Zhao, and Junchi Yan · 2023
Closest in time.