Fetching the paper…
Reading the bibliography…
A central goal of visual recognition is to understand objects and scenes from a single image.
Description and recognition of curved objects
Ramakant Nevatia and Thomas O Binford · 1977
Earlier work this paper cites.
Geometric modeling using octree encoding
Donald Meagher · 1982
Earlier work this paper cites.
Ray tracing volume densities
James T Kajiya and Brian P Von Herzen · 1984
Earlier work this paper cites.
Estimating uncertain spatial relationships in robotics
Randall Smith, Matthew Self, and Peter Cheeseman · 1990
Earlier work this paper cites.
Shape and motion from image streams under orthography: a factorization method
Carlo Tomasi and Takeo Kanade · 1992
Earlier work this paper cites.
The SPmap: A probabilistic framework for simultaneous localization and map building
Jose A Castellanos, José MM Montiel, José Neira, and Juan D Tardós · 1999
Earlier work this paper cites.
A taxonomy and evaluation of dense two-frame stereo correspondence algorithms
Daniel Scharstein and Richard Szeliski · 2002
Earlier work this paper cites.
Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman · 2003
Earlier work this paper cites.
Silhouette and stereo fusion for 3D object modeling
Carlos Hernández Esteban and Francis Schmitt · 2004
Earlier work this paper cites.
Nonrigid structure-from-motion: Estimating shape and motion with hierarchical priors
Lorenzo Torresani, Aaron Hertzmann, and Chris Bregler · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Acquiring 3D indoor environments with variability and repetition
Young Min Kim, Niloy J Mitra, Dong-Ming Yan, and Leonidas Guibas · 2012
Earlier work this paper cites.
A search-classify approach for cluttered indoor scene understanding
Liangliang Nan, Ke Xie, and Andrei Sharf · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
ShapeNet: An information-rich 3D model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Joint 3D object and layout inference from a single RGB-D image
Andreas Geiger and Chaohui Wang · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
RAPter: rebuilding man-made scenes with regular arrangements of planes
Aron Monszpart, Nicolas Mellado, Gabriel J Brostow, and Niloy J Mitra · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Structured prediction of unobserved voxels from a single depth image
Michael Firman, Oisin Mac Aodha, Simon Julier, and Gabriel J Brostow · 2016
Earlier work this paper cites.
Learning a predictable and generative vector representation for objects
Rohit Girdhar, David F Fouhey, Mikel Rodriguez, and Abhinav Gupta · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Pixelwise view selection for unstructured multi-view stereo
Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Conditional image generation with PixelCNN decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Perspective transformer nets: Learning single-view 3D object reconstruction without 3D supervision
Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, and Honglak Lee · 2016
Earlier work this paper cites.
A point set generation network for 3D object reconstruction from a single image
Haoqiang Fan, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
PointNet: Deep learning on point sets for 3D classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
MarrNet: 3D shape reconstruction via 2.5D sketches
Jiajun Wu, Yifan Wang, Tianfan Xue, Xingyuan Sun, Bill Freeman, and Josh Tenenbaum · 2017
Cited alongside, same era.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Cited alongside, same era.
Learning category-specific mesh reconstruction from image collections
Angjoo Kanazawa, Shubham Tulsiani, Alexei A Efros, and Jitendra Malik · 2018
Cited alongside, same era.
Neural 3D mesh renderer
Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada · 2018
Cited alongside, same era.
3D-RCNN: Instance-level 3D object reconstruction via render-and-compare
Abhijit Kundu, Yin Li, and James M. Rehg · 2018
Cited alongside, same era.
NeRF: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Later among the works it cites.
Differentiable volumetric rendering: Learning implicit 3D representations without 3D supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger · 2020
Later among the works it cites.
Accelerating 3D deep learning with PyTorch3D
Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari · 2020
Later among the works it cites.
Unsupervised learning of probably symmetric deformable 3D objects from images in the wild
Shangzhe Wu, Christian Rupprecht, and Andrea Vedaldi · 2020
Later among the works it cites.
NeRF++: Analyzing and improving neural radiance fields
Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pix3d: Dataset and methods for single-image 3d shape modeling
Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenenbaum, and William T Freeman · 2018
Cited alongside, same era.
Pixel2Mesh: Generating 3D mesh models from single RGB images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang · 2018
Cited alongside, same era.
PCN: Point completion network
Wentao Yuan, Tejas Khot, David Held, Christoph Mertz, and Martial Hebert · 2018
Cited alongside, same era.
Taskonomy: Disentangling task transfer learning
Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese · 2018
Cited alongside, same era.
Learning to predict 3D objects with an interpolation-based differentiable renderer
Wenzheng Chen, Huan Ling, Jun Gao, Edward Smith, Jaakko Lehtinen, Alec Jacobson, and Sanja Fidler · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Unsupervised learning of 3D object categories from videos in the wild
Philipp Henzler, Jeremy Reizenstein, Patrick Labatut, Roman Shapovalov, Tobias Ritschel, Andrea Vedaldi, and David Novotny · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Pixel-aligned volumetric avatars
Amit Raj, Michael Zollhofer, Tomas Simon, Jason Saragih, Shunsuke Saito, James Hays, and Stephen Lombardi · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
Common objects in 3D: Large-scale learning and evaluation of real-life 3D category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny · 2021
Later among the works it cites.
Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind · 2021
Later among the works it cites.
Multi-view 3d reconstruction with transformers
Dan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou, Tianyang Shi, Septimiu Salcudean, Z Jane Wang, and Rabab Ward · 2021
Later among the works it cites.
PixelNeRF: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2021
Later among the works it cites.
PoinTr: Diverse point cloud completion with geometry-aware transformers
Xumin Yu, Yongming Rao, Ziyi Wang, Zuyan Liu, Jiwen Lu, and Jie Zhou · 2021
Later among the works it cites.
3D shape generation and completion through point-voxel diffusion
Linqi Zhou, Yilun Du, and Jiajun Wu · 2021
Later among the works it cites.
BEiT: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2022
Later among the works it cites.
MonoScene: Monocular 3D semantic scene completion
Anh-Quan Cao and Raoul de Charette · 2022
Later among the works it cites.
Depth-supervised NeRF: Fewer views and faster training for free
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Perceiver IO: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2022
Later among the works it cites.
What’s behind the couch? directed ray distance functions for 3D scene reconstruction
Nilesh Kulkarni, Justin Johnson, and David F. Fouhey · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
Computer vision: algorithms and applications
Richard Szeliski · 2022
Later among the works it cites.
Multi-view mesh reconstruction with neural deferred shading
Markus Worchel, Rodrigo Diaz, Weiwen Hu, Oliver Schreer, Ingo Feldmann, and Peter Eisert · 2022
Later among the works it cites.
ShapeFormer: Transformer-based shape completion via sparse representation
Xingguang Yan, Liqiang Lin, Niloy J Mitra, Dani Lischinski, Daniel Cohen-Or, and Hui Huang · 2022
Later among the works it cites.