Fetching the paper…
Reading the bibliography…
A comprehensive semantic understanding of a scene is important for many applications - but in what space should diverse semantic information (e.g., objects, scene categories, material types, texture, etc.) be grounded and what should be its structure? Aspiring to have one unified structure that hosts diverse types of semantics, we follow the Scene Graph paradigm in 3D, generating a 3D Scene Graph.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando CN Pereira · 2001
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Beyond categories: The visual memex model for reasoning about object relationships
Tomasz Malisiewicz and Alyosha Efros · 2009
Earlier work this paper cites.
Building a database of 3d scenes from user annotations
Bryan C Russell and Antonio Torralba · 2009
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Philipp Krähenbühl and Vladlen Koltun · 2011
Earlier work this paper cites.
Understanding indoor scenes using 3d geometric phrases
Wongun Choi, Yu-Wei Chao, Caroline Pantofaru, and Silvio Savarese · 2013
Earlier work this paper cites.
Scene parsing by integrating function, geometry and appearance models
Yibiao Zhao and Song-Chun Zhu · 2013
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva · 2014
Earlier work this paper cites.
Reasoning about object affordances in a knowledge base representation
Yuke Zhu, Alireza Fathi, and Li Fei-Fei · 2014
Earlier work this paper cites.
Material recognition in the wild with the materials in context database
Sean Bell, Paul Upchurch, Noah Snavely, and Kavita Bala · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
3d semantic parsing of large-scale indoor spaces
Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Structural-rnn: Deep learning on spatio-temporal graphs
Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Cited alongside, same era.
Amodal instance segmentation
Ke Li and Jitendra Malik · 2016
Cited alongside, same era.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Joint 2d-3d-semantic data for indoor scene understanding
Iro Armeni, Sasha Sax, Amir R Zamir, and Silvio Savarese · 2017
Frustum pointnets for 3d object detection from rgb-d data
Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas · 2017
Later among the works it cites.
Segcloud: Semantic segmentation of 3d point clouds
Lyne Tchapmi, Christopher Choy, Iro Armeni, JunYoung Gwak, and Silvio Savarese · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Later among the works it cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua · 2017
Later among the works it cites.
Semantic amodal segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Annotating object instances with a polygon-rnn
Lluis Castrejon, Kaustav Kundu, Raquel Urtasun, and Sanja Fidler · 2017
Cited alongside, same era.
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Cited alongside, same era.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas A Funkhouser, and Matthias Nießner · 2017
Cited alongside, same era.
Blitznet: A real-time deep network for scene understanding
Nikita Dvornik, Konstantin Shmelkov, Julien Mairal, and Cordelia Schmid · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Yan Zhu, Yuandong Tian, Dimitris Metaxas, and Piotr Dollár · 2017
Later among the works it cites.
Efficient interactive annotation of segmentation datasets with polygon-rnn++
David Acuna, Huan Ling, Amlan Kar, and Sanja Fidler · 2018
Later among the works it cites.
Fluid annotation: a human-machine collaboration interface for full image annotation
Mykhaylo Andriluka, Jasper RR Uijlings, and Vittorio Ferrari · 2018
Later among the works it cites.
Segan: Segmenting and generating the invisible
Kiana Ehsani, Roozbeh Mottaghi, and Ali Farhadi · 2018
Later among the works it cites.
Holistic 3d scene parsing and reconstruction from a single rgb image
Siyuan Huang, Siyuan Qi, Yixin Zhu, Yinxue Xiao, Yuanlu Xu, and Song-Chun Zhu · 2018
Later among the works it cites.
Configurable 3d scene synthesis and 2d image rendering with per-pixel ground truth using stochastic grammars
Chenfanfu Jiang, Siyuan Qi, Yixin Zhu, Siyuan Huang, Jenny Lin, Lap-Fai Yu, Demetri Terzopoulos, and Song-Chun Zhu · 2018
Later among the works it cites.
A robust 3d-2d interactive tool for scene segmentation and annotation
Duc Thanh Nguyen, Binh-Son Hua, Lap-Fai Yu, and Sai-Kit Yeung · 2018
Later among the works it cites.
Learning human-object interactions by graph parsing neural networks
Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, and Song-Chun Zhu · 2018
Later among the works it cites.
A deep learning based behavioral approach to indoor autonomous navigation
Gabriel Sepulveda, Juan Carlos Niebles, and Alvaro Soto · 2018
Later among the works it cites.
Gibson env: Real-world perception for embodied agents
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Later among the works it cites.
https://github.com/facebookresearch/Detectron/blob/master/MODEL_ZOO.md
Detectron model zoo · 2019
Closest in time.
Supplementary material for: 3D Scene Graph: A structure for unified semantics, 3D space, and camera
Iro Armeni, Jerry He, JunYoung Gwak, Amir R Zamir, Martin Fischer, Jitendra Malik, and Silvio Savarese · 2019
Closest in time.