Fetching the paper…
Reading the bibliography…
Grounding language to visual relations is critical to various language-and-vision applications.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
S. K. Divvala, A. Farhadi, and C. Guestrin · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Language models for image captioning: The quirks and what works
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
RNN fisher vectors for action recognition and image annotation
G. Lev, G. Sadeh, B. Klein, and L. Wolf · 2016
Cited alongside, same era.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Order-embeddings of images and language
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Cited alongside, same era.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Cited alongside, same era.
Linking image and text with 2-way nets
A. Eisenschtat and L. Wolf · 2017
Cited alongside, same era.
Phrase localization and visual relationship detection with comprehensive image-language cues
B. A. Plummer, A. Mallya, C. M. Cervantes, J. Hockenmaier, and S. Lazebnik · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2017
Later among the works it cites.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Later among the works it cites.
Ppr-fcn: weakly supervised visual relation detection via parallel pairwise r-fcn
H. Zhang, Z. Kyaw, J. Yu, and S.-F. Chang · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler · 2017
Cited alongside, same era.
Instance-aware image and sentence matching with selective multimodal LSTM
Y. Huang, W. Wang, and L. Wang · 2017
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Cited alongside, same era.
Vip-cnn: Visual phrase guided convolutional neural network
Y. Li, W. Ouyang, X. Wang, and X. Tang · 2017
Cited alongside, same era.
Scene graph generation from objects, phrases and region captions
Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang · 2017
Cited alongside, same era.
Z. Zheng, L. Zheng, M. Garrett, Y. Yang, and Y.-D. Shen · 2017
Later among the works it cites.
Towards context-aware interaction recognition for visual relationship detection
B. Zhuang, L. Liu, C. Shen, and I. Reid · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and VQA
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
J. Gu, J. Cai, S. Joty, L. Niu, and G. Wang · 2018
Later among the works it cites.
Learning semantic concepts and order for image and sentence matching
Y. Huang, Q. Wu, and L. Wang · 2018
Later among the works it cites.
Illustrative language understanding: Large-scale visual grounding with image search
J. Kiros, W. Chan, and G. Hinton · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He · 2018
Later among the works it cites.
Graph r-cnn for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Later among the works it cites.
Scene graph reasoning with prior visual relationship for visual question answering
Z. Yang, J. Yu, C. Yang, Z. Qin, and Y. Hu · 2018
Later among the works it cites.
Exploring visual relationship for image captioning
T. Yao, Y. Pan, Y. Li, and T. Mei · 2018
Later among the works it cites.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Later among the works it cites.
Relational reasoning using prior knowledge for visual captioning
J. Hou, X. Wu, Y. Qi, W. Zhao, J. Luo, and Y. Jia · 2019
Closest in time.
Object-driven text-to-image synthesis via adversarial training
W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu, and J. Gao · 2019
Closest in time.
Rethinking visual relationships for high-level image understanding
Y. Liang, Y. Bai, W. Zhang, X. Qian, L. Zhu, and T. Mei · 2019
Closest in time.