Fetching the paper…
Reading the bibliography…
Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data.
Visual entailment: A novel task for fine-grained image understanding
Xie, N.; Lai, F.; Doran, D.; and Kadav, A. 2019 · 1901
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P.; Larochelle, H.; Bengio, Y.; and Manzagol, P.-A. 2008 · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. 2011 · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Suhr, A.; Zhou, S.; Zhang, A.; Zhang, I.; Bai, H.; and Artzi, Y. 2018 · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Yu, L.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; Bansal, M.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Wang, W.; Bao, H.; Dong, L.; and Wei, F. 2021 · 2021
Later among the works it cites.
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Zeng, Y.; Zhang, X.; and Li, H. 2021 · 2021
Later among the works it cites.
Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Andonian, A.; Chen, S.; and Hamid, R. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020 · 2020
Cited alongside, same era.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Li, G.; Duan, N.; Fang, Y.; Gong, M.; and Jiang, D. 2020 · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
Bao, H.; Dong, L.; and Wei, F. 2021 · 2021
Cited alongside, same era.
Peco: Perceptual codebook for bert pre-training of vision transformers
Dong, X.; Bao, J.; Zhang, T.; Chen, D.; Zhang, W.; Yuan, L.; Chen, D.; Wen, F.; and Yu, N. 2021 · 2021
Cited alongside, same era.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021 · 2021
Cited alongside, same era.
Closest in time.
Context autoencoder for self-supervised representation learning
Chen, X.; Ding, M.; Wang, X.; Xin, Y.; Mo, S.; Wang, Y.; Han, S.; Luo, P.; Zeng, G.; and Wang, J. 2022 · 2022
Closest in time.
Multimodal Masked Autoencoders Learn Transferable Representations
Geng, X.; Liu, H.; Lee, L.; Schuurams, D.; Levine, S.; and Abbeel, P. 2022 · 2022
Closest in time.
Training Vision-Language Transformers from Captions Alone
Gui, L.; Huang, Q.; Hauptmann, A.; Bisk, Y.; and Gao, J. 2022 · 2022
Closest in time.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022 · 2022
Closest in time.
mc-BEiT: Multi-choice Discretization for Image BERT Pre-training
Li, X.; Ge, Y.; Yi, K.; Hu, Z.; Shan, Y.; and Duan, L.-Y. 2022 · 2022
Closest in time.
Masked feature prediction for self-supervised visual pre-training
Wei, C.; Fan, H.; Xie, S.; Wu, C.-Y.; Yuille, A.; and Feichtenhofer, C. 2022 · 2022
Closest in time.
Simmim: A simple framework for masked image modeling
Xie, Z.; Zhang, Z.; Cao, Y.; Lin, Y.; Bao, J.; Yao, Z.; Dai, Q.; and Hu, H. 2022 · 2022
Closest in time.
Vision-Language Pre-Training with Triple Contrastive Learning
Yang, J.; Duan, J.; Tran, S.; Xu, Y.; Chanda, S.; Chen, L.; Zeng, B.; Chilimbi, T.; and Huang, J. 2022 · 2022
Closest in time.