Fetching the paper…
Reading the bibliography…
Building scalable vision-language models to learn from diverse, multimodal data remains an open challenge.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019 · 1908
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. L. 2011 · 2011
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q. V.; Hinton, G. E.; and Dean, J. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
A Corpus for Reasoning about Natural Language Grounded in Photographs
Suhr, A.; Zhou, S.; Zhang, A.; Zhang, I.; Bai, H.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
UNITER: UNiversal Image-TExt Representation Learning
Chen, Y.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Earlier work this paper cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2020 · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
Hazimeh, H.; Zhao, Z.; Chowdhery, A.; Sathiamoorthy, M.; Chen, Y.; Mazumder, R.; Hong, L.; and Chi, E. H. 2021 · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q. V.; Sung, Y.; Li, Z.; and Duerig, T. 2021 · 2021
Cited alongside, same era.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2021 · 2021
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Later among the works it cites.
Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
Mustafa, B.; Riquelme, C.; Puigcerver, J.; Jenatton, R.; and Houlsby, N. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Later among the works it cites.
BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
Peng, Z.; Dong, L.; Bao, H.; Ye, Q.; and Wei, F. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BASE Layers: Simplifying Training of Large, Sparse Models
Lewis, M.; Bhosale, S.; Dettmers, T.; Goyal, N.; and Zettlemoyer, L. 2021 · 2021
Cited alongside, same era.
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Li, J.; Selvaraju, R. R.; Gotmare, A.; Joty, S. R.; Xiong, C.; and Hoi, S. C. 2021 · 2021
Cited alongside, same era.
OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
Liu, J.; Zhu, X.; Liu, F.; Guo, L.; Zhao, Z.; Sun, M.; Wang, W.; Lu, H.; Zhou, S.; Zhang, J.; and Wang, J. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Zero-Shot Text-to-Image Generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
Riquelme, C.; Puigcerver, J.; Mustafa, B.; Neumann, M.; Jenatton, R.; Pinto, A. S.; Keysers, D.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; Liu, C.; Liu, M.; Liu, Z.; Lu, Y.; Shi, Y.; Wang, L.; Wang, J.; Xiao, B.; Xiao, Z.; Yang, J.; Zeng, M.; Zhou, L.; and Zhang, P. 2021 · 2021
Cited alongside, same era.
Yang, J.; Duan, J.; Tran, S.; Xu, Y.; Chanda, S.; Chen, L.; Zeng, B.; Chilimbi, T.; and Huang, J. 2022 · 2022
Later among the works it cites.
FILIP: Fine-grained Interactive Language-Image Pre-Training
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2022 · 2022
Later among the works it cites.
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Zeng, Y.; Zhang, X.; and Li, H. 2022 · 2022
Later among the works it cites.
Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs
Zhu, J.; Zhu, X.; Wang, W.; Wang, X.; Li, H.; Wang, X.; and Dai, J. 2022 · 2022
Later among the works it cites.
ST-MoE: Designing Stable and Transferable Sparse Expert Models
Zoph, B.; Bello, I.; Kumar, S.; Du, N.; Huang, Y.; Dean, J.; Shazeer, N.; and Fedus, W. 2022 · 2022
Later among the works it cites.
Write and Paint: Generative Vision-Language Models are Unified Modal Learners
Diao, S.; Zhou, W.; Zhang, X.; and Wang, J. 2023 · 2023
Closest in time.
Training Vision-Language Transformers from Captions
Gui, L.; Chang, Y.; Huang, Q.; Som, S.; Hauptmann, A.; Gao, J.; and Bisk, Y. 2023 · 2023
Closest in time.
Masked Vision and Language Modeling for Multi-modal Representation Learning
Kwon, G.; Cai, Z.; Ravichandran, A.; Bas, E.; Bhotika, R.; and Soatto, S. 2023 · 2023
Closest in time.
Prismer: A Vision-Language Model with An Ensemble of Experts
Liu, S.; Fan, L.; Johns, E.; Yu, Z.; Xiao, C.; and Anandkumar, A. 2023 · 2023
Closest in time.
Scaling Vision-Language Models with Sparse Mixture of Experts
Shen, S.; Yao, Z.; Li, C.; Darrell, T.; Keutzer, K.; and He, Y. 2023 · 2023
Closest in time.
Image as a foreign language: BEiT pretraining for vision and vision-language tasks
Wang, W.; Bao, H.; Dong, L.; Bjorck, J.; Peng, Z.; Liu, Q.; Aggarwal, K.; Mohammed, O. K.; Singhal, S.; Som, S.; and Wei, F. 2023 · 2023
Closest in time.
Zhang, X.; Zeng, Y.; Zhang, J.; and Li, H. 2023 · 2023
Closest in time.