Fetching the paper…
Reading the bibliography…
We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language.
Ernie 2.0: A continual pre-training framework for language understanding
Sun, Y.; Wang, S.; Li, Y.; Feng, S.; Tian, H.; Wu, H.; and Wang, H. 2019 · 1907
Earlier work this paper cites.
An empirical study on leveraging scene graphs for visual question answering
Zhang, C.; Chao, W.-L.; and Xuan, D. 2019 · 1907
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Li, G.; Duan, N.; Fang, Y.; Jiang, D.; and Zhou, M. 2019a · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H.; Yatskar, M.; Yin, D.; Hsieh, C.-J.; and Chang, K.-W. 2019b · 1908
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H.; and Bansal, M. 2019 · 1908
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Chen, Y.-C.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2019 · 1909
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Zhou, L.; Palangi, H.; Zhang, L.; Hu, H.; Corso, J. J.; and Gao, J. 2019 · 1909
Earlier work this paper cites.
Circle loss: A unified perspective of pair similarity optimization
Sun, Y.; Cheng, C.; Zhang, Y.; Zhang, C.; Zheng, L.; Wang, Z.; and Wei, Y. 2020 · 2002
Earlier work this paper cites.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Huang, Z.; Zeng, Z.; Liu, B.; Fu, D.; and Fu, J. 2020 · 2004
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2004
Earlier work this paper cites.
Large-Scale Adversarial Training for Vision-and-Language Representation Learning
Gan, Z.; Chen, Y.-C.; Li, L.; Zhu, C.; Cheng, Y.; and Liu, J. 2020 · 2006
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. L. 2011 · 2011
Cited alongside, same era.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Image generation from scene graphs
Johnson, J.; Gupta, A.; and Fei-Fei, L. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Cited alongside, same era.
Image retrieval using scene graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016 · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Yu, L.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; Bansal, M.; and Berg, T. L. 2018 · 2018
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Later among the works it cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Wu, H.; Mao, J.; Zhang, Y.; Jiang, Y.; Li, L.; Sun, W.; and Ma, W.-Y. 2019 · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Later among the works it cites.