Fetching the paper…
Reading the bibliography…
Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks.
VisualBERT: A Simple and Performant Baseline for Vision and Language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.; and Chang, K. 2019 · 1908
Earlier work this paper cites.
Unifying Vision-and-Language Tasks via Text Generation
Cho, J.; Lei, J.; Tan, H.; and Bansal, M. 2021 · 1942
Earlier work this paper cites.
Scene Graph Generation With External Knowledge and Image Reconstruction
Gu, J.; Zhao, H.; Lin, Z.; Li, S.; Cai, J.; and Ling, M. 2019 · 1978
Earlier work this paper cites.
InterBERT: Vision-and-Language Interaction for Multi-modal Pretraining
Lin, J.; Yang, A.; Zhang, Y.; Liu, J.; Zhou, J.; and Yang, H. 2020 · 2003
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Scene Graph Generation by Iterative Message Passing
Xu, D.; Zhu, Y.; Choy, C. B.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Faghri, F.; Fleet, D. J.; Kiros, J. R.; and Fidler, S. 2018 · 2018
Earlier work this paper cites.
Image Generation From Scene Graphs
Johnson, J.; Gupta, A.; and Fei-Fei, L. 2018 · 2018
Earlier work this paper cites.
Graph R-CNN for Scene Graph Generation
Yang, J.; Lu, J.; Lee, S.; Batra, D.; and Parikh, D. 2018 · 2018
Earlier work this paper cites.
Neural Motifs: Scene Graph Parsing With Global Context
Zellers, R.; Yatskar, M.; Thomson, S.; and Choi, Y. 2018 · 2018
Earlier work this paper cites.
Knowledge-Embedded Routing Network for Scene Graph Generation
Chen, T.; Yu, W.; Chen, R.; and Lin, L. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Cited alongside, same era.
Learning to Compose Dynamic Tree Structures for Visual Contexts
Tang, K.; Zhang, H.; Wu, B.; Luo, W.; and Liu, W. 2019 · 2019
Cited alongside, same era.
Auto-Encoding Scene Graphs for Image Captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Cited alongside, same era.
An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
Zhang, C.; Chao, W.; and Xuan, D. 2019 · 2019
Cited alongside, same era.
UNITER: UNiversal Image-TExt Representation Learning
Chen, Y.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs
Yu, F.; Tang, J.; Yin, W.; Sun, Y.; Tian, H.; Wu, H.; and Wang, H. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Yuan, W.; Neubig, G.; and Liu, P. 2021 · 2021
Later among the works it cites.
VinVL: Revisiting Visual Representations in Vision-Language Models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Later among the works it cites.
Aspect-based Sentiment Classification with Sequential Cross-modal Semantic Graph
Huang, Y.; Chen, Z.; Zhang, W.; Chen, J.; Pan, J. Z.; Yao, Z.; Xie, Y.; and Chen, H. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training
Li, G.; Duan, N.; Fang, Y.; Gong, M.; and Jiang, D. 2020 · 2020
Cited alongside, same era.
OntoZSL: Ontology-enhanced Zero-shot Learning
Geng, Y.; Chen, J.; Chen, Z.; Pan, J. Z.; Ye, Z.; Yuan, Z.; Jia, Y.; and Chen, H. 2021 · 2021
Cited alongside, same era.
Seeing Out of the Box: End-to-End Pre-Training for Vision-Language Representation Learning
Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021 · 2021
Cited alongside, same era.
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Huo, Y.; Zhang, M.; Liu, G.; Lu, H.; Gao, Y.; Yang, G.; Wen, J.; Zhang, H.; Xu, B.; Zheng, W.; Xi, Z.; Yang, Y.; Hu, A.; Zhao, J.; Li, R.; Zhao, Y.; Zhang, L.; Song, Y.; Hong, X.; Cui, W.; Hou, D. Y.; Li, Y.; Li, J.; Liu, P.; Gong, Z.; Jin, C.; Sun, Y.; Chen, S.; Lu, Z.; Dou, Z.; Jin, Q.; Lan, Y.; Zhao, W. X.; Song, R.; and Wen, J. 2021 · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q. V.; Sung, Y.; Li, Z.; and Duerig, T. 2021 · 2021
Cited alongside, same era.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. C. H. 2022 · 2022
Later among the works it cites.
FLAVA: A Foundational Language And Vision Alignment Model
Singh, A.; Hu, R.; Goswami, V.; Couairon, G.; Galuba, W.; Rohrbach, M.; and Kiela, D. 2022 · 2022
Later among the works it cites.
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
Thrush, T.; Jiang, R.; Bartolo, M.; Singh, A.; Williams, A.; Kiela, D.; and Ross, C. 2022 · 2022
Later among the works it cites.
CoCa: Contrastive Captioners are Image-Text Foundation Models
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022 · 2022
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Yüksekgönül, M.; Bianchi, F.; Kalluri, P.; Jurafsky, D.; and Zou, J. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. ???? · 2022
Later among the works it cites.
VisualGPTScore: Visio-Linguistic Reasoning with Multimodal Generative Pre-Training Scores
Lin, Z.; Chen, X.; Pathak, D.; Zhang, P.; and Ramanan, D. 2023 · 2023
Closest in time.
Zhang, Y.; Chen, Z.; and Zhang, W. 2023 · 2023
Closest in time.