Fetching the paper…
Reading the bibliography…
Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language.
Representation of manipulable man-made objects in the dorsal stream
L. L. Chao and A. Martin · 2000
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. S. Davis · 2008
Earlier work this paper cites.
Visualizing data using t-sne
L. v. d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
A. Gupta, A. Kembhavi, and L. S. Davis · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
Efficient object category recognition using classemes
L. Torresani, M. Szummer, and A. Fitzgibbon · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Discriminative models for multi-class object layout
C. Desai, D. Ramanan, and C. C. Fowlkes · 2011
Earlier work this paper cites.
Recognition using visual phrases
M. A. Sadeghi and A. Farhadi · 2011
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Label-embedding for image classification
Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Question answering over freebase with multi-column convolutional neural networks
L. Dong, F. Wei, M. Zhou, and K. Xu · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, D. J. Rezende, and D. Wierstra · 2015
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Later among the works it cites.
Learning to generalize to new compositions in image understanding
Y. Atzmon, J. Berant, V. Kezami, A. Globerson, and G. Chechik · 2016
Later among the works it cites.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Spatial transformer networks
M. Jaderberg, K. Simonyan, A. Zisserman, et al · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Learning entity and relation embeddings for knowledge graph completion
Y. Lin, Z. Liu, M. Sun, Y. Liu, and X. Zhu · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Learning semantic relationships for better action retrieval in images
V. Ramanathan, C. Li, J. Deng, W. Han, Z. Li, K. Gu, Y. Song, S. Bengio, C. Rossenberg, and L. Fei-Fei · 2015
Cited alongside, same era.
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Later among the works it cites.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, and S. Reed · 2016
Later among the works it cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Learning models for actions and person-object interactions with transfer to question answering
A. Mallya and S. Lazebnik · 2016
Later among the works it cites.
A review of relational machine learning for knowledge graphs
M. Nickel, K. Murphy, V. Tresp, and E. Gabrilovich · 2016
Later among the works it cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Later among the works it cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2016
Later among the works it cites.
Fvqa: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. v. d. Hengel, and A. Dick · 2016
Later among the works it cites.