Fetching the paper…
Reading the bibliography…
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Gpu clusters for high-performance computing
V. V. Kindratenko, J. J. Enos, G. Shi, M. T. Showerman, G. W. Arnold, J. E. Stone, J. C. Phillips, and W.-m. Hwu · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Multi-gpu mapreduce on gpu clusters
J. A. Stuart and J. D. Owens · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Earlier work this paper cites.
Adding chinese captions to images
X. Li, W. Lan, J. Dong, and H. Liu · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Earlier work this paper cites.
Image pivoting for learning multilingual multimodal representations
S. Gella, R. Sennrich, F. Keller, and M. Lapata · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Fluency-guided cross-lingual image captioning
W. Lan, X. Li, and J. Dong · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
Ai challenger: A large-scale dataset for going deeper in image understanding
J. Wu, H. Zheng, B. Zhao, Y. Li, B. Yan, R. Liang, W. Wang, S. Zhou, G. Lin, Y. Fu, et al · 2017
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. Van Der Maaten · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Directional skip-gram: Explicitly distinguishing left and right context for word embeddings
Y. Song, S. Shi, J. Li, and H. Zhang · 2018
Cited alongside, same era.
Learning two-branch neural networks for image-text matching tasks
L. Wang, Y. Li, J. Huang, and S. Lazebnik · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Cited alongside, same era.
Wenlan 2.0: Make ai imagine via a multimodal foundation model
N. Fei, Z. Lu, Y. Gao, G. Yang, Y. Huo, J. Wen, H. Lu, R. Song, X. Gao, T. Xiang, et al · 2021
Later among the works it cites.
Wenlan: Bridging vision and language by large-scale multi-modal pre-training
Y. Huo, M. Zhang, G. Liu, H. Lu, Y. Gao, G. Yang, J. Wen, H. Zhang, B. Xu, W. Zheng, et al · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
W. Kim, B. Son, and I. Kim · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Li, C. Xu, X. Wang, W. Lan, Z. Jia, G. Yang, and J. Xu · 2019
Cited alongside, same era.
Vilbert: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
How multilingual is multilingual bert?
T. Pires, E. Schlinger, and D. Garrette · 2019
Cited alongside, same era.
Language-agnostic visual-semantic embeddings
J. Wehrmann, D. M. Souza, M. A. Lopes, and R. C. Barros · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia, et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. A. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, et al · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
M. Ni, H. Huang, L. Su, E. Cui, T. Bharti, L. Wang, D. Zhang, and N. Duan · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby · 2021
Later among the works it cites.
Tokenlearner: Adaptive space-time tokenization for videos
M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova · 2021
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork · 2021
Later among the works it cites.
Lightningdot: Pre-training visual-semantic embeddings for real-time image-text retrieval
S. Sun, Y.-C. Chen, L. Li, S. Wang, Y. Fang, and J. Liu · 2021
Later among the works it cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2021
Later among the works it cites.
Product1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining
X. Zhan, Y. Wu, X. Dong, Y. Wei, M. Lu, Y. Zhang, H. Xu, and X. Liang · 2021
Later among the works it cites.
Uc2: Universal cross-lingual cross-modal vision-and-language pre-training
M. Zhou, L. Zhou, S. Wang, Y. Cheng, L. Li, Z. Yu, and J. Liu · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu · 2022
Closest in time.