Fetching the paper…
Reading the bibliography…
Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks.
Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning . 1931–1942
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021 · 1942
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011 · 2011
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. 2016 · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Fluency-guided cross-lingual image captioning. In Proceedings of the 25th ACM international conference on Multimedia . 1549–1557
Weiyu Lan, Xirong Li, and Jianfeng Dong. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives. In Proceedings of the British Machine Vision Conference
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Human Language Technology Conference of the NAACL . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 2019
Earlier work this paper cites.
COCO-CN for cross-lingual image tagging, captioning, and retrieval
Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, and Jieping Xu. 2019a · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Ai challenger: A large-scale dataset for going deeper in image understanding. In IEEE International Conference on Multimedia and Expo
Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, et al · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning. In European Conference on Computer Vision . 104–120
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Revisiting Pre-Trained Models for Chinese Natural Language Processing. In Conference on Empirical Methods in Natural Language Processing . 657–668
M6: A chinese multimodal pretrainer. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3251–3261
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2443–2449
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, and Guoping Hu. 2020 · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9729–9738
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Cited alongside, same era.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In Proceedings of the AAAI Conference on Artificial Intelligence . 11336–11344
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020 · 2020
Cited alongside, same era.
12-in-1: Multi-task vision and language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10437–10446
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2020 · 2020
Cited alongside, same era.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. 2020 · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3558–3568
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Cited alongside, same era.
RedCaps: Web-curated image-text data created by the people, for the people
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson. 2021 · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Cited alongside, same era.
Aligning pretraining for detection via object-level contrastive learning
Fangyun Wei, Yue Gao, Zhirong Wu, Han Hu, and Stephen Lin. 2021 · 2021
Later among the works it cites.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, et al · 2022
Closest in time.
Wukong: 100 Million Large-scale Chinese Cross-modal Pre-training Dataset and A Foundation Framework
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Minzhe Niu, Hang Xu, Xiaodan Liang, Wei Zhang, Xin Jiang, and Chunjing Xu. 2022 · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Closest in time.
FLAVA: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15638–15650
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022 · 2022
Closest in time.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning . 23318–23340
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Closest in time.
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
An Yang, Junshu Pan, Junyang Lin, Rui Men, Yichang Zhang, Jingren Zhou, and Chang Zhou. 2022 · 2022
Closest in time.
FILIP: Fine-grained Interactive Language-Image Pre-Training. In International Conference on Learning Representations
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2022 · 2022
Closest in time.