Fetching the paper…
Reading the bibliography…
Existing vision-language pre-training (VLP) methods primarily rely on paired image-text datasets, which are either annotated by enormous human labors, or crawled from the internet followed by elaborate data cleaning techniques.
Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training
Li, G., Duan, N., Fang, Y., Jiang, D., and Zhou, M · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 1908
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. B · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T · 2011
Earlier work this paper cites.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., and Wang, H · 2012
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
DeVries, T. and Taylor, G. W · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Unpaired image captioning by language pivoting
Gu, J., Joty, S., Cai, J., and Wang, G · 2018
Cited alongside, same era.
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Duerig, T., et al · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.-F., and Wang, W. Y · 2019
Later among the works it cites.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Wei, J. and Zou, K · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Yu, Z., Yu, J., Cui, Y., Tao, D., and Tian, Q · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y · 2018
Cited alongside, same era.
Uniter: Learning universal image-text representations
Chen, Y.-C., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Unsupervised image captioning
Feng, Y., Ma, L., Liu, W., and Luo, J · 2019
Cited alongside, same era.
Unpaired image captioning via scene graph alignments
Gu, J., Joty, S., Cai, J., Zhao, H., Yang, X., and Wang, G · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Cited alongside, same era.
VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Cao, J., Gan, Z., Cheng, Y., Yu, L., Chen, Y.-C., and Liu, J · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Later among the works it cites.
Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training
Hu, X., Yin, X., Lin, K., Wang, L., Zhang, L., Gao, J., and Liu, Z · 2020
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B., and Kim, I · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Later among the works it cites.