Fetching the paper…
Reading the bibliography…
Relationships among objects play a crucial role in image understanding.
Probabilistic reasoning in intelligent systems: Networks of plausible reasoning, 1988
Judea Pearl · 1988
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando CN Pereira · 2001
Earlier work this paper cites.
Measuring the similarity of labeled graphs
Pierre-Antoine Champin and Christine Solnon · 2003
Earlier work this paper cites.
Conditional random fields for object recognition
Ariadna Quattoni, Michael Collins, and Trevor Darrell · 2004
Earlier work this paper cites.
Discovering objects and their location in images
Josef Sivic, Bryan C Russell, Alexei A Efros, Andrew Zisserman, and William T Freeman · 2005
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
Bryan C Russell, William T Freeman, Alexei A Efros, Josef Sivic, and Andrew Zisserman · 2006
Earlier work this paper cites.
Objects in context
Andrew Rabinovich, Andrea Vedaldi, Carolina Galleguillos, Eric Wiewiora, and Serge Belongie · 2007
Earlier work this paper cites.
Towards scalable representations of object categories: Learning a hierarchy of parts
Sanja Fidler and Ales Leonardis · 2007
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
Abhinav Gupta and Larry S Davis · 2008
Earlier work this paper cites.
Object categorization using co-occurrence, location and appearance
Carolina Galleguillos, Andrew Rabinovich, and Serge Belongie · 2008
Earlier work this paper cites.
Putting objects in perspective
Derek Hoiem, Alexei A Efros, and Martial Hebert · 2008
Earlier work this paper cites.
Multi-class segmentation with relative location prior
Stephen Gould, Jim Rodgers, David Cohen, Gal Elidan, and Daphne Koller · 2008
Earlier work this paper cites.
Grouplet: A structured image representation for recognizing human and object interactions
Bangpeng Yao and Li Fei-Fei · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Context based object categorization: A critical survey
Carolina Galleguillos and Serge Belongie · 2010
Earlier work this paper cites.
Efficiently selecting regions for scene understanding
M Pawan Kumar and Daphne Koller · 2010
Earlier work this paper cites.
Exploiting hierarchical context on a large database of object categories
Myung Jin Choi, Joseph J Lim, Antonio Torralba, and Alan S Willsky · 2010
Earlier work this paper cites.
Graph cut based inference with co-occurrence statistics
Lubor Ladicky, Chris Russell, Pushmeet Kohli, and Philip HS Torr · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
Recognition using visual phrases
Mohammad Amin Sadeghi and Ali Farhadi · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C Berg, and Tamara L Berg · 2011
Earlier work this paper cites.
Learning to share visual appearance for multiclass object detection
Ruslan Salakhutdinov, Antonio Torralba, and Josh Tenenbaum · 2011
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Vladlen Koltun · 2011
Cited alongside, same era.
Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation
Jian Yao, Sanja Fidler, and Raquel Urtasun · 2012
Cited alongside, same era.
Understanding and predicting importance in images
Alexander C Berg, Tamara L Berg, Hal Daume, Jesse Dodge, Amit Goyal, Xufeng Han, Alyssa Mensch, Margaret Mitchell, Aneesh Sood, Karl Stratos, et al · 2012
Cited alongside, same era.
Understanding indoor scenes using 3d geometric phrases
Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, and Silvio Savarese · 2013
Cited alongside, same era.
Image description using visual dependency representations
Desmond Elliott and Frank Keller · 2013
Cited alongside, same era.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A Shamma, Michael S Bernstein, and Li Fei-Fei · 2015
Later among the works it cites.
Contextual action recognition with r* cnn
Georgia Gkioxari, Ross Girshick, and Jitendra Malik · 2015
Later among the works it cites.
Learning semantic relationships for better action retrieval in images
Vignesh Ramanathan, Congcong Li, Jia Deng, Wei Han, Zhen Li, Kunlong Gu, Yang Song, Samy Bengio, Chuck Rossenberg, and Li Fei-Fei · 2015
Later among the works it cites.
Sherlock: Scalable fact learning in images
Mohamed Elhoseiny, Scott Cohen, Walter Chang, Brian Price, and Ahmed Elgammal · 2015
Later among the works it cites.
Recognize complex events from static images by fusing deep channels
Yuanjun Xiong, Kai Zhu, Dahua Lin, and Xiaoou Tang · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Translating video content to natural language descriptions
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Cited alongside, same era.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2013
Cited alongside, same era.
Learning the visual interpretation of sentences
C Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende · 2013
Cited alongside, same era.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Cited alongside, same era.
Panda: Pose aligned networks for deep attribute modeling
Ning Zhang, Manohar Paluri, Marc’Aurelio Ranzato, Trevor Darrell, and Lubomir Bourdev · 2014
Cited alongside, same era.
Integrating language and vision to generate natural language descriptions of videos in the wild
Jesse Thomason, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Raymond J Mooney · 2014
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2015
Later among the works it cites.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Later among the works it cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Conditional random fields as recurrent neural networks
Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip HS Torr · 2015
Later among the works it cites.
Learning deep structured models
Liang-Chieh Chen, Alexander G Schwing, Alan L Yuille, and Raquel Urtasun · 2015
Later among the works it cites.
Fully connected deep structured networks
Alexander G Schwing and Raquel Urtasun · 2015
Later among the works it cites.
Structured prediction energy networks
David Belanger and Andrew McCallum · 2015
Later among the works it cites.
From images to sentences through scene description graphs using commonsense reasoning and knowledge
Somak Aditya, Yezhou Yang, Chitta Baral, Cornelia Fermuller, and Yiannis Aloimonos · 2015
Later among the works it cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Later among the works it cites.
Places: An image database for deep scene understanding
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Antonio Torralba, and Aude Oliva · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Later among the works it cites.
Deep markov random field for image modeling
Zhirong Wu, Dahua Lin, and Xiaoou Tang · 2016
Later among the works it cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2016
Later among the works it cites.