Fetching the paper…
Reading the bibliography…
Existing text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query).
Does the wake-sleep algorithm produce good density estimators?
Brendan J Frey, Geoffrey E Hinton, and Peter Dayan · 1995
Earlier work this paper cites.
Information theory and statistics
Solomon Kullback · 1997
Earlier work this paper cites.
Distance metric learning with application to clustering with side-information
Eric Xing, Michael Jordan, Stuart J Russell, and Andrew Ng · 2002
Earlier work this paper cites.
Generative or discriminative? getting the best of both worlds
JM Bernardo, MJ Bayarri, JO Berger, AP Dawid, D Heckerman, AFM Smith, and M West · 2007
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William Dolan · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Earlier work this paper cites.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Earlier work this paper cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Earlier work this paper cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Cited alongside, same era.
Segdiff: Image segmentation with diffusion probabilistic models
Tomer Amit, Eliya Nachmani, Tal Shaharbany, and Lior Wolf · 2021
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg · 2021
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Cited alongside, same era.
Dynamic clustering network for unsupervised semantic segmentation
Kehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu, Yian Zhao, Guoli Song, Li Yuan, and Jie Chen · 2022
Later among the works it cites.
Locality guidance for improving vision transformers on tiny datasets
Kehan Li, Runyi Yu, Zhennan Wang, Li Yuan, Guoli Song, and Jie Chen · 2022
Later among the works it cites.
Neighborhood-adaptive structure augmented metric learning
Pandeng Li, Yan Li, Hongtao Xie, and Lei Zhang · 2022
Later among the works it cites.
Dual-stream knowledge-preserving hashing for unsupervised video retrieval
Pandeng Li, Hongtao Xie, Jiannan Ge, Lei Zhang, Shaobo Min, and Yongdong Zhang · 2022
Later among the works it cites.
Deep fourier ranking quantization for semi-supervised image retrieval
Pandeng Li, Hongtao Xie, Shaobo Min, Jiannan Ge, Xun Chen, and Yongdong Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dmitry Baranchuk, Ivan Rubachev, Andrey Voynov, Valentin Khrulkov, and Artem Babenko · 2021
Cited alongside, same era.
Teachtext: Crossmodal generalized distillation for text-video retrieval
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Cited alongside, same era.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu · 2021
Cited alongside, same era.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G Hauptmann, Joao F. Henriques, and Andrea Vedaldi · 2021
Cited alongside, same era.
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B Hashimoto · 2022
Later among the works it cites.
Gmmseg: Gaussian mixture based generative semantic segmentation models
Chen Liang, Wenguan Wang, Jiaxu Miao, and Yi Yang · 2022
Later among the works it cites.
Ts2-net: Token shift and selection transformer for text-video retrieval
Yuqi Liu, Pengfei Xiong, Luhui Xu, Shengming Cao, and Qin Jin · 2022
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano · 2022
Later among the works it cites.
Disentangled representation learning for text-video retrieval
Qiang Wang, Yanhao Zhang, Yun Zheng, Pan Pan, and Xian-Sheng Hua · 2022
Later among the works it cites.
Zero-shot image restoration using denoising diffusion null-space model
Yinhuai Wang, Jiwen Yu, and Jian Zhang · 2022
Later among the works it cites.
Position embedding needs an independent layer normalization
Runyi Yu, Zhennan Wang, Yinhuai Wang, Kehan Li, Yian Zhao, Jian Zhang, Guoli Song, and Jie Chen · 2022
Later among the works it cites.
Align representations with base: A new approach to self-supervised learning
Shaofeng Zhang, Lyn Qiu, Feng Zhu, Junchi Yan, Hengrui Zhang, Rui Zhao, Hongyang Li, and Xiaokang Yang · 2022
Later among the works it cites.
Tagging before alignment: Integrating multi-modal tags for video-text retrieval
Yizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi, Jin Ma, and Ying Shan · 2023
Closest in time.
Wico: Win-win cooperation of bottom-up and top-down referring image segmentation
Zesen Cheng, Peng Jin, Hao Li, Kehan Li, Siheng Li, Xiangyang Ji, Chang Liu, and Jie Chen · 2023
Closest in time.
Parallel vertex diffusion for unified visual grounding
Zesen Cheng, Kehan Li, Peng Jin, Xiangyang Ji, Li Yuan, Chang Liu, and Jie Chen · 2023
Closest in time.
Uatvr: Uncertainty-adaptive text-video retrieval
Bo Fang, Chang Liu, Yu Zhou, Min Yang, Yuxin Song, Fu Li, Weiping Wang, Xiangyang Ji, Wanli Ouyang, et al · 2023
Closest in time.
Video-text as game players: Hierarchical banzhaf interaction for cross-modal representation learning
Peng Jin, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu, Xiangyang Ji, Li Yuan, and Jie Chen · 2023
Closest in time.
Text-video retrieval with disentangled conceptualization and set-to-set alignment
Peng Jin, Hao Li, Zesen Cheng, Jinfa Huang, Zhennan Wang, Li Yuan, Chang Liu, and Jie Chen · 2023
Closest in time.
Weakly-supervised 3d spatial reasoning for text-based visual question answering
Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen · 2023
Closest in time.
Tg-vqa: Ternary game of video question answering
Hao Li, Peng Jin, Zesen Cheng, Songyang Zhang, Kai Chen, Zhennan Wang, Chang Liu, and Jie Chen · 2023
Closest in time.
Multi-granularity interaction simulation for unsupervised interactive segmentation
Kehan Li, Yian Zhao, Zhennan Wang, Zesen Cheng, Peng Jin, Xiangyang Ji, Li Yuan, Chang Liu, and Jie Chen · 2023
Closest in time.
Momentdiff: Generative video moment retrieval from random to real
Pandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao, Lei Zhang, Yun Zheng, Deli Zhao, and Yongdong Zhang · 2023
Closest in time.
Shaofeng Zhang, Feng Zhu, Rui Zhao, and Junchi Yan · 2023
Closest in time.