Fetching the paper…
Reading the bibliography…
We integrate two powerful ideas, geometry and deep visual representation learning, into recurrent network architectures for mobile visual scene understanding.
Perception of partly occluded objects by young chicks
L. Regolin and G. Vallortigara · 1995
Earlier work this paper cites.
C. E. Freer, D. M. Roy, and J. B. Tenenbaum · 2012
Earlier work this paper cites.
Dense visual SLAM for RGB-D cameras
C. Kerl, J. Sturm, and D. Cremers · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
D. Eigen, C. Puhrsch, and R. Fergus · 2014
Earlier work this paper cites.
Semi-dense visual odometry for AR on a smartphone
T. Schöps, J. Engel, and D. Cremers · 2014
Earlier work this paper cites.
P. Agrawal, J. Carreira, and J. Malik · 2015
Earlier work this paper cites.
ShapeNet: An Information-Rich 3D Model Repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, J. Xiao, L. Yi, and F. Yu · 2015
Earlier work this paper cites.
Learning image representations equivariant to ego-motion
D. Jayaraman and K. Grauman · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Single-view to multi-view: Reconstructing unseen views with a convolutional network
M. Tatarchenko, A. Dosovitskiy, and T. Brox · 2015
Cited alongside, same era.
Unsupervised monocular depth estimation with left-right consistency
C. Godard, O. Mac Aodha, and G. J. Brostow · 2016
Cited alongside, same era.
IQA: visual question answering in interactive environments
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi · 2017
Cited alongside, same era.
Cognitive mapping and planning for visual navigation
S. Gupta, J. Davidson, S. Levine, R. Sukthankar, and J. Malik · 2017
Cited alongside, same era.
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Marrnet: 3d shape reconstruction via 2.5d sketches
J. Wu, Y. Wang, T. Xue, X. Sun, W. T. Freeman, and J. B. Tenenbaum · 2017
Later among the works it cites.
Da-rnn: Semantic mapping with data associated recurrent neural networks
Y. Xiang and D. Fox · 2017
Later among the works it cites.
Unsupervised learning of depth and ego-motion from video
T. Zhou, M. Brown, N. Snavely, and D. G. Lowe · 2017
Later among the works it cites.
Geometry-aware recurrent neural networks for active visual recognition
R. Cheng, Z. Wang, and K. Fragkiadaki · 2018
Closest in time.
Neural scene representation and rendering
S. M. A. Eslami, D. Jimenez Rezende, F. Besse, F. Viola, A. S. Morcos, M. Garnelo, A. Ruderman, A. A. Rusu, I. Danihelka, K. Gregor, D. P. Reichert, L. Buesing, T. Weber, O. Vinyals, D. Rosenbaum, N. Rabinowitz, H. King, C. Hillier, M. Botvinick, D. Wierstra, K. Kavukcuoglu, and D. Hassabis · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning a multi-view stereo machine
A. Kar, C. Häne, and J. Malik · 2017
Cited alongside, same era.
Learning 3d object categories by looking around them
D. Novotný, D. Larlus, and A. Vedaldi · 2017
Cited alongside, same era.
Neural map: Structured memory for deep reinforcement learning
E. Parisotto and R. Salakhutdinov · 2017
Cited alongside, same era.
Multi-view supervision for single-view reconstruction via differentiable ray consistency
S. Tulsiani, T. Zhou, A. A. Efros, and J. Malik · 2017
Cited alongside, same era.
Adversarial inverse graphics networks: Learning 2d-to-3d lifting and image-to-image translation with unpaired supervision
H. F. Tung, A. Harley, W. Seto, and K. Fragkiadaki · 2017
Cited alongside, same era.
Sfm-net: Learning of structure and motion from video
S. Vijayanarasimhan, S. Ricco, C. Schmid, R. Sukthankar, and K. Fragkiadaki · 2017
Cited alongside, same era.
Mapnet: An allocentric spatial memory for mapping environments
J. F. Henriques and A. Vedaldi · 2018
Closest in time.
Fusion++: Volumetric object-level slam
M. B. A. J. D. John McCormac, Ronald Clark and S. Leutenegger · 2018
Closest in time.
Deep continuous fusion for multi-sensor 3d object detection
M. Liang, B. Yang, S. Wang, and R. Urtasun · 2018
Closest in time.
3d interpreter networks for viewer-centered wireframe modeling
J. Wu, T. Xue, J. J. Lim, Y. Tian, J. B. Tenenbaum, A. Torralba, and W. T. Freeman · 2018
Closest in time.
Hdnet: Exploiting hd maps for 3d object detection
B. Yang, M. Liang, and R. Urtasun · 2018
Closest in time.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Y. Zhou and O. Tuzel · 2018
Closest in time.