Fetching the paper…
Reading the bibliography…
We propose a transformer-based neural network architecture for multi-object 3D reconstruction from RGB videos.
Pollefeys, M., Koch, R., Van Gool, L.: Self-calibration and metric reconstruction inspite of varying and unknown intrinsic camera parameters. IJCV
1999
Earlier work this paper cites.
Lewiner, T., Lopes, H., Vieira, A.W., Tavares, G.: Efficient implementation of marching cubes’ cases with topological guarantees. J. Graphics, GPU, & Game Tools
2003
Earlier work this paper cites.
Frome, A., Huber, D., Kolluri, R., Bülow, T., Malik, J.: Recognizing objects in range data using regional point descriptors. In: ECCV (2004)
2004
Earlier work this paper cites.
Andriluka, M., Roth, S., Schiele, B.: People-tracking-by-detection and people-detection-by-tracking. In: CVPR (2008)
2008
Earlier work this paper cites.
Breitenstein, M.D., Reichlin, F., Leibe, B., Koller-Meier, E., Van Gool, L.: Robust tracking-by-detection using a detector confidence particle filter. In: ICCV (2009)
2009
Earlier work this paper cites.
Nan, L., Xie, K., Sharf, A.: A search-classify approach for cluttered indoor scene understanding. ACM Transactions on Graphics (TOG) (2012)
2012
Earlier work this paper cites.
Shao, T., Xu, W., Zhou, K., Wang, J., Li, D., Guo, B.: An interactive approach to semantic modeling of indoor scenes with an RGBD camera. ACM Transactions on Graphics (TOG) (2012)
2012
Earlier work this paper cites.
Salas-Moreno, R.F., Newcombe, R.A., Strasdat, H., Kelly, P.H., Davison, A.J.: SLAM++: Simultaneous localisation and mapping at the level of objects. In: CVPR (2013)
2013
Earlier work this paper cites.
Wu, C.: Towards linear-time incremental structure from motion. In: 3DV (2013)
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
Li, Y., Dai, A., Guibas, L., Nießner, M.: Database-assisted object retrieval for real-time 3d reconstruction. In: Computer Graphics Forum. vol. 34. Wiley Online Library (2015)
2015
Earlier work this paper cites.
Mur-Artal, R., Montiel, J.M.M., Tardos, J.D.: ORB-SLAM: a versatile and accurate monocular slam system. IEEE transactions on robotics (2015)
2015
Earlier work this paper cites.
Choy, C.B., Xu, D., Gwak, J., Chen, K., Savarese, S.: 3D-R2N2: A unified approach for single and multi-view 3D object reconstruction. In: ECCV (2016)
2016
Earlier work this paper cites.
Girdhar, R., Fouhey, D., Rodriguez, M., Gupta, A.: Learning a predictable and generative vector representation for objects. In: ECCV (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
2016
Earlier work this paper cites.
Schönberger, J.L., Frahm, J.M.: Structure-from-motion revisited. In: CVPR (2016)
2016
Earlier work this paper cites.
Wu, J., Zhang, C., Xue, T., Freeman, W.T., Tenenbaum, J.B.: Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In: NIPS (2016)
2016
Earlier work this paper cites.
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Nießner, M.: Scannet: Richly-annotated 3d reconstructions of indoor scenes. In: CVPR (2017)
2017
Earlier work this paper cites.
Engel, J., Koltun, V., Cremers, D.: Direct sparse odometry. TPAMI
2017
Earlier work this paper cites.
Izadinia, H., Shan, Q., Seitz, S.M.: Im2CAD. In: CVPR (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Mousavian, A., Anguelov, D., Flynn, J., Kosecka, J.: 3d bounding box estimation using deep learning and geometry. In: CVPR (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: NeurIPS (2017)
2017
Cited alongside, same era.
Fei, X., Soatto, S.: Visual-inertial object detection and mapping. In: ECCV (2018)
2018
Cited alongside, same era.
Huang, S., Qi, S., Zhu, Y., Xiao, Y., Xu, Y., Zhu, S.C.: Holistic 3D scene parsing and reconstruction from a single RGB image. In: ECCV (2018)
2018
Cited alongside, same era.
Kundu, A., Li, Y., Rehg, J.M.: 3D-RCNN: Instance-level 3d object reconstruction via render-and-compare. In: CVPR (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Kuo, W., Angelova, A., Lin, T.Y., Dai, A.: Mask2CAD: 3D shape prediction by learning to segment and retrieve. In: ECCV (2020)
2020
Later among the works it cites.
Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. In: ECCV (2020)
2020
Later among the works it cites.
Nie, Y., Han, X., Guo, S., Zheng, Y., Chang, J., Zhang, J.J.: Total3dunderstanding: Joint layout, object pose and mesh reconstruction for indoor scenes from a single image. In: CVPR (2020)
2020
Later among the works it cites.
Popov, S., Bauszat, P., Ferrari, V.: CoReNet: Coherent 3D scene reconstruction from a single RGB image. In: ECCV (2020)
2020
Later among the works it cites.
Qian, S., Jin, L., Fouhey, D.F.: Associative3d: Volumetric reconstruction from sparse views. In: ECCV (2020)
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholson, L., Milford, M., Sünderhauf, N.: Quadricslam: Dual quadrics from object detections as landmarks in object-oriented slam. RA-L
2018
Cited alongside, same era.
Tulsiani, S., Gupta, S., Fouhey, D., Efros, A.A., Malik, J.: Factoring shape, pose, and layout from the 2d image of a 3d scene. In: CVPR (2018)
2018
Cited alongside, same era.
Wang, N., Zhang, Y., Li, Z., Fu, Y., Liu, W., Jiang, Y.G.: Pixel2mesh: Generating 3d mesh models from single RGB images. In: ECCV (2018)
2018
Cited alongside, same era.
Avetisyan, A., Dahnert, M., Dai, A., Savva, M., Chang, A.X., Nießner, M.: Scan2CAD: Learning cad model alignment in RGB-D scans. In: CVPR (2019)
2019
Cited alongside, same era.
Avetisyan, A., Dai, A., Nießner, M.: End-to-end cad model retrieval and 9dof alignment in 3d scans. In: ICCV (2019)
2019
Cited alongside, same era.
Bergmann, P., Meinhardt, T., Leal-Taixe, L.: Tracking without bells and whistles. In: ICCV (2019)
2019
Cited alongside, same era.
Gkioxari, G., Malik, J., Johnson, J.: Mesh R-CNN. In: ICCV (2019)
2019
Cited alongside, same era.
Later among the works it cites.
Runz, M., Li, K., Tang, M., Ma, L., Kong, C., Schmidt, T., Reid, I., Agapito, L., Straub, J., Lovegrove, S., et al.: Frodo: From detections to 3d objects. In: CVPR (2020)
2020
Later among the works it cites.
Xie, H., Yao, H., Zhang, S., Zhou, S., Sun, W.: Pix2vox++: multi-scale context-aware 3d object reconstruction from single and multiple images. IJCV
2020
Later among the works it cites.
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C.: ViViT: A video vision transformer. In: CVPR (2021)
2021
Later among the works it cites.
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? In: ICML (2021)
2021
Later among the works it cites.
Cheng, B., Schwing, A.G., Kirillov, A.: Per-pixel classification is not all you need for semantic segmentation. In: NeurIPS (2021)
2021
Later among the works it cites.
Duzceker, A., Galliani, S., Vogel, C., Speciale, P., Dusmanu, M., Pollefeys, M.: Deepvideomvs: Multi-view stereo on video with recurrent spatio-temporal fusion. In: CVPR (2021)
2021
Later among the works it cites.
Engelmann, F., Rematas, K., Leibe, B., Ferrari, V.: From points to multi-object 3d reconstruction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4588–4597 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Kuo, W., Angelova, A., Lin, T.Y., Dai, A.: Patch2cad: Patchwise embedding learning for in-the-wild shape retrieval from a single image. In: ICCV (2021)
2021
Later among the works it cites.
Li, K., DeTone, D., Chen, Y.F.S., Vo, M., Reid, I., Rezatofighi, H., Sweeney, C., Straub, J., Newcombe, R.: Odam: Object detection, association, and mapping using posed rgb video. In: ICCV (2021)
2021
Later among the works it cites.
Li, K., Rezatofighi, H., Reid, I.: MOLTR: Multiple object localization, tracking and reconstruction from monocular rgb videos. IEEE Robotics and Automation Letters
2021
Later among the works it cites.
Meinhardt, T., Kirillov, A., Leal-Taixe, L., Feichtenhofer, C.: Trackformer: Multi-object tracking with transformers. arXiv (2021)
2021
Later among the works it cites.
Shan, M., Feng, Q., Jau, Y.Y., Atanasov, N.: ELLIPSDF: Joint object pose and shape optimization with a bi-level ellipsoid and signed distance function description. In: ICCV (2021)
2021
Later among the works it cites.
Strudel, R., Garcia, R., Laptev, I., Schmid, C.: Segmenter: Transformer for semantic segmentation. In: ICCV (2021)
2021
Later among the works it cites.
Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P.H., et al.: Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: CVPR (2021)
2021
Later among the works it cites.
Danila Rukhovich, Anna Vorontsova, A.K.: ImVoxelNet: Image to voxels projection for monocular and multi-view general-purpose 3d object detection. In: WACV (2022)
2022
Closest in time.
Maninis, K.K., Popov, S., Niesser, M., Ferrari, V.: Vid2CAD: Cad model alignment using multi-view constraints from videos. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
Closest in time.