WenLan: Bridging vision and language by large-scale multi-modal pre-training
Original
Yuqi Huo, Manli Zhang, Guangzhen Liu, Haoyu Lu, Yizhao Gao, Guoxing Yang, Jingyuan Wen, Heng Zhang, Baogui Xu, Weihao Zheng, et al · 2021
Later among the works it cites.
Improving image captioning by leveraging intra-and inter-layer global representation in transformer network. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 1655–1663
Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu, Yue Gao, and Rongrong Ji. 2021 · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning . PMLR, 4904–4916
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Later among the works it cites.
Hierarchical Cross-Modal Graph Consistency Learning for Video-Text Retrieval. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021 , Fernando Diaz, Chirag Shah, Torsten Suel, Pablo Castells, Rosie Jones, and Tetsuya Sakai (Eds.). ACM, 1114–1124
Weike Jin, Zhou Zhao, Pengcheng Zhang, Jieming Zhu, Xiuqiang He, and Yueting Zhuang. 2021 · 2021
Later among the works it cites.
Relevance-guided supervision for openqa with colbert
Omar Khattab, Christopher Potts, and Matei Zaharia. 2021 · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7331–7341
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021 · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021 · 2021
Later among the works it cites.
Hit: Hierarchical transformer with momentum contrast for video-text retrieval
Original
Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, and Zhongyuan Wang. 2021 · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Original
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021 · 2021
Later among the works it cites.
A straightforward framework for video retrieval using clip. In Mexican Conference on Pattern Recognition . Springer, 3–12
Jesús Andrés Portillo-Quintero, José Carlos Ortiz-Bayliss, and Hugo Terashima-Marín. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Original
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Colbertv2: Effective and efficient retrieval via lightweight late interaction
Original
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021 · 2021
Later among the works it cites.
T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision Foundation / IEEE, 5079–5088
Xiaohan Wang, Linchao Zhu, and Yi Yang. 2021 · 2021
Later among the works it cites.
E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, Online, 503–513
Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang. 2021 · 2021
Later among the works it cites.
Taco: Token-aware cascade contrastive learning for video-text alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11562–11572
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
FILIP: Fine-grained Interactive Language-Image Pre-Training
Original
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning with adaptive attention on visual and non-visual words. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 15465–15474
Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021 · 2021
Later among the works it cites.
PixelFolder: An Efficient Progressive Pixel Synthesis Network for Image Generation
Original
Jing He, Yiyi Zhou, Qi Zhang, Jun Peng, Yunhang Shen, Xiaoshuai Sun, Chao Chen, and Rongrong Ji. 2022 · 2022
Closest in time.
Knowing What to Learn: A Metric-oriented Focal Mechanism for Image Captioning
Jiayi Ji, Yiwei Ma, Xiaoshuai Sun, Yiyi Zhou, Yongjian Wu, and Rongrong Ji. 2022 · 2022
Closest in time.
Knowing what it is: Semantic-enhanced Dual Attention Transformer
Yiwei Ma, Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2022 · 2022
Closest in time.
SeqTR: A Simple yet Universal Network for Visual Grounding
Original
Chaoyang Zhu, Yiyi Zhou, Yunhang Shen, Gen Luo, Xingjia Pan, Mingbao Lin, Chao Chen, Liujuan Cao, Xiaoshuai Sun, and Rongrong Ji. 2022 · 2022
Closest in time.