Fetching the paper…
Reading the bibliography…
Vision-Language Pretraining (VLP) models have recently successfully facilitated many cross-modal downstream tasks.
HAKE: Human Activity Knowledge Engine
Li, Y.-L.; Xu, L.; Huang, X.; Liu, X.; Ma, Z.; Chen, M.; Wang, S.; Fang, H.; and Lu, C. 2019b · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019a · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H.; and Bansal, M. 2019 · 1908
Earlier work this paper cites.
Grounded Situation Recognition
Pratt, S.; Yatskar, M.; Weihs, L.; Farhadi, A.; and Kembhavi, A. 2020 · 2003
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Ribeiro, M. T.; Wu, T.; Guestrin, C.; and Singh, S. 2020 · 2005
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Cited alongside, same era.
Tide: A general toolbox for identifying object detection errors
Bolya, D.; Foley, S.; Hays, J.; and Hoffman, J. 2020 · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Later among the works it cites.
VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words
Lu, X.; Zhao, T.; and Lee, K. 2021 · 2021
Later among the works it cites.
Learning to Predict Visual Attributes in the Wild
Pham, K.; Kafle, K.; Lin, Z.; Ding, Z.; Cohen, S. D.; Tran, Q.; and Shrivastava, A. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Cited alongside, same era.
Decoupling the role of data, attention, and losses in multimodal transformers
Hendricks, L. A.; Mellor, J.; Schneider, R.; Alayrac, J.-B.; and Nematzadeh, A. 2021 · 2021
Cited alongside, same era.
Probing image-language transformers for verb understanding
Hendricks, L. A.; and Nematzadeh, A. 2021 · 2021
Cited alongside, same era.
MDETR-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Kiela, D.; Bartolo, M.; Nie, Y.; Kaushik, D.; Geiger, A.; Wu, Z.; Vidgen, B.; Prasad, G.; Singh, A.; Ringshia, P.; et al. 2021 · 2021
Cited alongside, same era.
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Dou, Z.-Y.; Xu, Y.; Gan, Z.; Wang, J.; Wang, S.; Wang, L.; Zhu, C.; Zhang, P.; Yuan, L.; Peng, N.; et al. 2022 · 2022
Closest in time.
Du, X.; Legastelois, B.; Ganesh, B.; Rajan, A.; Chockler, H.; Belle, V.; Anderson, S.; and Ramamoorthy, S. 2022 · 2022
Closest in time.
ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension
Subramanian, S.; Merrill, W.; Darrell, T.; Gardner, M.; Singh, S.; and Rohrbach, A. 2022 · 2022
Closest in time.
All in one: Exploring unified video-language pre-training
Wang, A. J.; Ge, Y.; Yan, R.; Ge, Y.; Lin, X.; Cai, G.; Wu, J.; Shan, Y.; Qie, X.; and Shou, M. Z. 2022 · 2022
Closest in time.
Vision-Language Pre-Training with Triple Contrastive Learning
Yang, J.; Duan, J.; Tran, S.; Xu, Y.; Chanda, S.; Chen, L.; Zeng, B.; Chilimbi, T.; and Huang, J. 2022 · 2022
Closest in time.
Visual Commonsense in Pretrained Unimodal and Multimodal Models
Zhang, C.; Van Durme, B.; Li, Z.; and Stengel-Eskin, E. 2022 · 2022
Closest in time.