Fetching the paper…
Reading the bibliography…
Answering complex questions about images is an ambitious goal for machine intelligence, which requires a joint understanding of images, text, and commonsense knowledge, as well as a strong reasoning ability.
Simple bert models for relation extraction and semantic role labeling
Shi, P.; and Lin, J. 2019 · 1904
Earlier work this paper cites.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Huang, Z.; Zeng, Z.; Liu, B.; Fu, D.; and Fu, J. 2020 · 2004
Earlier work this paper cites.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Li, X.; Yin, X.; Li, C.; Hu, X.; Zhang, P.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020b · 2004
Earlier work this paper cites.
GATE: Graph Attention Transformer Encoder for Cross-lingual Relation and Event Extraction
Ahmad, W. U.; Peng, N.; and Chang, K.-W. 2020 · 2010
Earlier work this paper cites.
Weakly-supervised VisualBERT: Pre-training without Parallel Images and Captions
Li, L. H.; You, H.; Wang, Z.; Zareian, A.; Chang, S.-F.; and Chang, K.-W. 2020a · 2010
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Image Retrieval Using Scene Graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. 2016 · 2016
Earlier work this paper cites.
Yin and Yang: Balancing and Answering Binary Visual Questions
Zhang, P.; Goyal, Y.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
End-to-end neural coreference resolution
Lee, K.; He, L.; Lewis, M.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Scene Graph Generation by Iterative Message Passing
Xu, D.; Zhu, Y.; Choy, C.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Image Generation from Scene Graphs
Johnson, J.; Gupta, A.; and Fei-Fei, L. 2018 · 2018
Cited alongside, same era.
Referring Relationships
Krishna, R.; Chami, I.; Bernstein, M.; and Fei-Fei, L. 2018 · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training
Radford, A.; and Narasimhan, K. 2018 · 2018
Cited alongside, same era.
Visual entailment task for visually-grounded language learning
Xie, N.; Lai, F.; Doran, D.; and Kadav, A. 2018 · 2018
Cited alongside, same era.
Graph r-cnn for scene graph generation
Yang, J.; Lu, J.; Lee, S.; Batra, D.; and Parikh, D. 2018 · 2018
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Later among the works it cites.
A Corpus for Reasoning about Natural Language Grounded in Photographs
Suhr, A.; Zhou, S.; Zhang, A.; Zhang, I.; Bai, H.; and Artzi, Y. 2019 · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Later among the works it cites.
Auto-Encoding Scene Graphs for Image Captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Later among the works it cites.
From Recognition to Cognition: Visual Commonsense Reasoning
Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning
Chen, Y.-C.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural Motifs: Scene Graph Parsing with Global Context
Zellers, R.; Yatskar, M.; Thomson, S.; and Choi, Y. 2018 · 2018
Cited alongside, same era.
Knowledge-Embedded Routing Network for Scene Graph Generation
Chen, T.; Yu, W.; Chen, R.; and Lin, L. 2019 · 2019
Cited alongside, same era.
Class-balanced loss based on effective number of samples
Cui, Y.; Jia, M.; Lin, T.-Y.; Song, Y.; and Belongie, S. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Cited alongside, same era.
Explainable and Explicit Visual Reasoning over Scene Graphs
Jiaxin Shi, J. L., Hanwang Zhang. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Large-Scale Adversarial Training for Vision-and-Language Representation Learning
Gan, Z.; Chen, Y.-C.; Li, L.; Zhu, C.; Cheng, Y.; and Liu, J. 2020 · 2020
Later among the works it cites.
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
Jiang, X.; Yu, J.; Qin, Z.; Zhuang, Y.; Zhang, X.; Hu, Y.; and Wu, Q. 2020 · 2020
Later among the works it cites.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2020 · 2020
Later among the works it cites.
ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Yu, F.; Tang, J.; Yin, W.; Sun, Y.; Tian, H.; Wu, H.; and Wang, H. 2020 · 2020
Later among the works it cites.
Weakly Supervised Visual Semantic Parsing
Zareian, A.; Karaman, S.; and Chang, S.-F. 2020 · 2020
Later among the works it cites.
Learning Visual Commonsense for Robust Scene Graph Generation
Zareian, A.; You, H.; Wang, Z.; and Chang, S.-F. 2020 · 2020
Later among the works it cites.
Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
Zhu, Z.; Yu, J.; Sun, Y.; Hu, Y.; Wang, Y.; and Wu, Q. 2020 · 2020
Later among the works it cites.