Fetching the paper…
Reading the bibliography…
3D scene representation for robot manipulation should capture three key object properties: permanency -- objects that become occluded over time continue to exist; amodal completeness -- objects have 3D occupancy, even if only partial observations are available; spatiotemporal continuity -- the movement of each object is continuous over space and time.
Active perception
R. Bajcsy · 1988
Earlier work this paper cites.
Core knowledge
E. S. Spelke and K. D. Kinzler · 2007
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. Van Der Smagt, D. Cremers, and T. Brox · 2015
Earlier work this paper cites.
Unsupervised cnn for single view depth estimation: Geometry to the rescue
R. Garg, V. K. BG, G. Carneiro, and I. Reid · 2016
Earlier work this paper cites.
Self-supervised visual descriptor learning for dense correspondence
T. Schmidt, R. Newcombe, and D. Fox · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Earlier work this paper cites.
Sfm-net: Learning of structure and motion from video
S. Vijayanarasimhan, S. Ricco, C. Schmid, R. Sukthankar, and K. Fragkiadaki · 2017
Earlier work this paper cites.
Se3-nets: Learning rigid body motion using deep neural networks
A. Byravan and D. Fox · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser · 2017
Earlier work this paper cites.
Shape completion enabled robotic grasping
J. Varley, C. DeChant, A. Richardson, J. Ruales, and P. Allen · 2017
Earlier work this paper cites.
Learning a multi-view stereo machine
A. Kar, C. Häne, and J. Malik · 2017
Earlier work this paper cites.
3dmatch: Learning local geometric descriptors from rgb-d reconstructions
A. Zeng, S. Song, M. Nießner, M. Fisher, J. Xiao, and T. Funkhouser · 2017
Earlier work this paper cites.
Interactive perception: Leveraging action in perception and perception in action
J. Bohg, K. Hausman, B. Sankaran, O. Brock, D. Kragic, S. Schaal, and G. S. Sukhatme · 2017
Cited alongside, same era.
Se3-pose-nets: Structured deep dynamics models for visuomotor planning and control
A. Byravan, F. Leeb, F. Meier, and D. Fox · 2018
Cited alongside, same era.
Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction
H. Zhan, R. Garg, C. Saroj Weerasekera, K. Li, H. Agarwal, and I. Reid · 2018
Cited alongside, same era.
Revisiting active perception
R. Bajcsy, Y. Aloimonos, and J. K. Tsotsos · 2018
Cited alongside, same era.
Deepmvs: Learning multi-view stereopsis
P.-H. Huang, K. Matzen, J. Kopf, N. Ahuja, and J.-B. Huang · 2018
Cited alongside, same era.
Dense object nets: Learning dense visual object descriptors by and for robotic manipulation
Learning to infer implicit surfaces without 3d supervision
S. Liu, S. Saito, W. Chen, and H. Li · 2019
Later among the works it cites.
Object discovery in videos as foreground motion clustering
C. Xie, Y. Xiang, Z. Harchaoui, and D. Fox · 2019
Later among the works it cites.
Learning visual dynamics models of rigid objects using relational inductive biases
F. Ferreira, L. Shao, T. Asfour, and J. Bohg · 2019
Later among the works it cites.
Object-centric forward modeling for model predictive control
Y. Ye, D. Gandhi, A. Gupta, and S. Tulsiani · 2019
Later among the works it cites.
Reasoning about physical interactions with object-oriented prediction and planning
M. Janner, S. Levine, W. T. Freeman, J. B. Tenenbaum, C. Finn, and J. Wu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. R. Florence, L. Manuelli, and R. Tedrake · 2018
Cited alongside, same era.
3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation
A. Dai and M. Nießner · 2018
Cited alongside, same era.
Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans
A. Dai, D. Ritchie, M. Bokeloh, S. Reed, J. Sturm, and M. Nießner · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Cited alongside, same era.
Occupancy networks: Learning 3d reconstruction in function space
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger · 2019
Cited alongside, same era.
Learning implicit fields for generative shape modeling
Z. Chen and H. Zhang · 2019
Cited alongside, same era.
Z. Xu, J. Wu, A. Zeng, J. B. Tenenbaum, and S. Song · 2019
Later among the works it cites.
Monet: Unsupervised scene decomposition and representation
C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner · 2019
Later among the works it cites.
Embodied language grounding with 3d visual feature representations
M. Prabhudesai, H.-Y. F. Tung, S. A. Javed, M. Sieb, A. W. Harley, and K. Fragkiadaki · 2020
Closest in time.
Local implicit grid representations for 3d scenes
C. M. Jiang, A. Sud, A. Makadia, J. Huang, M. Niessner, and T. Funkhouser · 2020
Closest in time.
Learning continuous 3d reconstructions for geometrically aware grasping
M. Van der Merwe, Q. Lu, B. Sundaralingam, M. Matak, and T. Hermans · 2020
Closest in time.
Learning from unlabelled videos using contrastive predictive neural 3d mapping
A. W. Harley, S. K. Lakshmikanth, F. Li, X. Zhou, H.-Y. F. Tung, and K. Fragkiadaki · 2020
Closest in time.
Spatial action maps for mobile manipulation
J. Wu, X. Sun, A. Zeng, S. Song, J. Lee, S. Rusinkiewicz, and T. Funkhouser · 2020
Closest in time.