Fetching the paper…
Reading the bibliography…
We present a Deep Cuboid Detector which takes a consumer-quality RGB image of a cluttered scene and localizes all 3D cuboids (box-like objects).
Machine perception of three-dimensional soups
L. G. Roberts · 1963
Earlier work this paper cites.
Recognition-by-components: a theory of human image understanding
I. Biederman · 1987
Earlier work this paper cites.
Multiple view geometry in computer vision
R. Hartley and A. Zisserman · 2003
Earlier work this paper cites.
Using geometric constraints through parallelepipeds for calibration and 3d modeling
M. Wilczkowiak, P. Sturm, and E. Boyer · 2005
Earlier work this paper cites.
3d generic object categorization, localization and pose estimation
S. Savarese and L. Fei-Fei · 2007
Earlier work this paper cites.
Geometric reasoning for single image structure recovery
D. C. Lee, M. Hebert, and T. Kanade · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Blocks world revisited: Image understanding using qualitative geometry and mechanics
A. Gupta, A. A. Efros, and M. Hebert · 2010
Earlier work this paper cites.
The moped framework: Object recognition and pose estimation for manipulation
A. Collet, M. Martinez, and S. S. Srinivasa · 2011
Earlier work this paper cites.
Joint 3d estimation of objects and scene layout
A. Geiger, C. Wojek, and R. Urtasun · 2011
Earlier work this paper cites.
From 3d scene geometry to human workspace
A. Gupta, S. Satkin, A. A. Efros, and M. Hebert · 2011
Earlier work this paper cites.
Representations and techniques for 3d object recognition and scene interpretation
D. Hoiem and S. Savarese · 2011
Earlier work this paper cites.
3d object detection and viewpoint estimation with a deformable 3d cuboid model
S. Fidler, S. Dickinson, and R. Urtasun · 2012
Earlier work this paper cites.
Recovering free space of indoor scenes from a single image
V. Hedau, D. Hoiem, and D. Forsyth · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Localizing 3d cuboids in single-view images
J. Xiao, B. Russell, and A. Torralba · 2012
Earlier work this paper cites.
Interactive images: cuboid proxies for smart image manipulation
Y. Zheng, X. Chen, M.-M. Cheng, K. Zhou, S.-M. Hu, and N. J. Mitra · 2012
Earlier work this paper cites.
Data-driven 3d primitives for single image understanding
D. F. Fouhey, A. Gupta, and M. Hebert · 2013
Earlier work this paper cites.
3d-based reasoning with blocks, support, and stability
Z. Jia, A. Gallagher, A. Saxena, and T. Chen · 2013
Earlier work this paper cites.
A linear approach to matching cuboids in rgbd images
H. Jiang and J. Xiao · 2013
Cited alongside, same era.
Articulated human detection with flexible mixtures of parts
Y. Yang and D. Ramanan · 2013
Cited alongside, same era.
Seeing 3d chairs: exemplar part-based 2d-3d alignment using a large dataset of cad models
M. Aubry, D. Maturana, A. A. Efros, B. C. Russell, and J. Sivic · 2014
Cited alongside, same era.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Cited alongside, same era.
Spatial pyramid pooling in deep convolutional networks for visual recognition
Viewpoints and keypoints
S. Tulsiani and J. Malik · 2015
Later among the works it cites.
Data-driven 3d voxel patterns for object category recognition
Y. Xiang, W. Choi, Y. Lin, and S. Savarese · 2015
Later among the works it cites.
Marr revisited: 2d-3d alignment via surface normal prediction
A. Bansal, B. Russell, and A. Gupta · 2016
Closest in time.
Recurrent human pose estimation
V. Belagiannis and A. Zisserman · 2016
Closest in time.
Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks
S. Bell, C. L. Zitnick, K. Bala, and R. Girshick · 2016
Closest in time.
Human pose estimation via convolutional part heatmap regression
A. Bulat and G. Tzimiropoulos · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Fpm: Fine pose parts-based model with 3d cad models
J. J. Lim, A. Khosla, and A. Torralba · 2014
Cited alongside, same era.
Imagining the unseen: Stability-based cuboid arrangements for scene understanding
T. Shao, A. Monszpart, Y. Zheng, B. Koo, W. Xu, K. Zhou, and N. J. Mitra · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sliding shapes for 3d object detection in depth images
S. Song and J. Xiao · 2014
Cited alongside, same era.
A novel representation of parts for accurate 3d object detection and tracking in monocular images
A. Crivellaro, M. Rad, Y. Verdie, K. Moo Yi, P. Fua, and V. Lepetit · 2015
Cited alongside, same era.
Closest in time.
Human pose estimation with iterative error feedback
J. Carreira, P. Agrawal, K. Fragkiadaki, and J. Malik · 2016
Closest in time.
3d-r2n2: A unified approach for single and multi-view 3d object reconstruction
C. B. Choy, D. Xu, J. Gwak, K. Chen, and S. Savarese · 2016
Closest in time.
Instance-aware semantic segmentation via multi-task network cascades
J. Dai, K. He, and J. Sun · 2016
Closest in time.
Deep image homography estimation
D. DeTone, T. Malisiewicz, and A. Rabinovich · 2016
Closest in time.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
S. Han, H. Mao, and W. J. Dally · 2016
Closest in time.
Categorizing cubes: Revisiting pose normalization
M. Hejrati and D. Ramanan · 2016
Closest in time.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 1mb model size
F. N. Iandola, M. W. Moskewicz, K. Ashraf, S. Han, W. J. Dally, and K. Keutzer · 2016
Closest in time.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, and S. Reed · 2016
Closest in time.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Closest in time.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Closest in time.
Deep sliding shapes for amodal 3d object detection in rgb-d images
S. Song and J. Xiao · 2016
Closest in time.
Single image 3d interpreter network
J. Wu, T. Xue, J. J. Lim, Y. Tian, J. B. Tenenbaum, A. Torralba, and W. T. Freeman · 2016
Closest in time.