Fetching the paper…
Reading the bibliography…
Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc.
The PASCAL visual object classes (VOC) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Detect what you can: Detecting and representing objects using holistic models and body parts
Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
3D menagerie: Modeling the 3D shape and pose of animals
Silvia Zuffi, Angjoo Kanazawa, David W Jacobs, and Michael J Black · 2017
Earlier work this paper cites.
Canonical surface mapping via geometric cycle consistency
Nilesh Kulkarni, Abhinav Gupta, and Shubham Tulsiani · 2019
Earlier work this paper cites.
Soft rasterizer: A differentiable renderer for image-based 3D reasoning
Shichen Liu, Tianye Li, Weikai Chen, and Hao Li · 2019
Earlier work this paper cites.
Three-D Safari: Learning to estimate zebra pose, shape, and texture from images "in the wild"
Silvia Zuffi, Angjoo Kanazawa, Tanya Berger-Wolf, and Michael J Black · 2019
Earlier work this paper cites.
Shape and viewpoint without keypoints
Shubham Goel, Angjoo Kanazawa, and Jitendra Malik · 2020
Earlier work this paper cites.
Articulation-aware canonical surface mapping
Nilesh Kulkarni, Abhinav Gupta, David F Fouhey, and Shubham Tulsiani · 2020
Earlier work this paper cites.
Online adaptation for consistent mesh reconstruction in the wild
Xueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, and Jan Kautz · 2020
Earlier work this paper cites.
Self-supervised single-view 3D reconstruction via semantic consistency
Xueting Li, Sifei Liu, Kihwan Kim, Shalini De Mello, Varun Jampani, Ming-Hsuan Yang, and Jan Kautz · 2020
Earlier work this paper cites.
NeRF: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Earlier work this paper cites.
Implicit mesh reconstruction from unannotated image collections
Shubham Tulsiani, Nilesh Kulkarni, and Abhinav Gupta · 2020
Earlier work this paper cites.
Deep ViT features as dense visual descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel · 2021
Earlier work this paper cites.
NeRD: Neural reflectance decomposition from image collections
Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Barron, Ce Liu, and Hendrik Lensch · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Pixel-aligned volumetric avatars
Amit Raj, Michael Zollhofer, Tomas Simon, Jason Saragih, Shunsuke Saito, James Hays, and Stephen Lombardi · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Ibrnet: Learning multi-view image-based rendering
Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srinivasan, Howard Zhou, Jonathan T Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser · 2021
Cited alongside, same era.
Latent-nerf for shape-guided generation of 3d shapes and textures
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or · 2022
Later among the works it cites.
Share with thy neighbors: Single-view reconstruction by cross-instance consistency
Tom Monnier, Matthew Fisher, Alexei A Efros, and Mathieu Aubry · 2022
Later among the works it cites.
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Huiwen Chang, Deva Ramanan, William T Freeman, and Ce Liu · 2021
Cited alongside, same era.
ViSER: Video-specific surface embeddings for articulated 3d shape reconstruction
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Ce Liu, and Deva Ramanan · 2021
Cited alongside, same era.
BANMo: Building animatable 3D neural models from many casual videos
Gengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan, Andrea Vedaldi, and Hanbyul Joo · 2021
Cited alongside, same era.
Shelf-supervised mesh prediction in the wild
Yufei Ye, Shubham Tulsiani, and Abhinav Gupta · 2021
Cited alongside, same era.
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2021
Cited alongside, same era.
NeRS: Neural reflectance surfaces for sparse-view 3D reconstruction in the wild
Jason Zhang, Gengshan Yang, Shubham Tulsiani, and Deva Ramanan · 2021
Cited alongside, same era.
PhySG: Inverse rendering with spherical Gaussians for physics-based material editing and relighting
Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely · 2021
Cited alongside, same era.
Later among the works it cites.
Sparf: Neural radiance fields from sparse and noisy poses
Prune Truong, Marie-Julie Rakotosaona, Fabian Manhardt, and Federico Tombari · 2022
Later among the works it cites.
Magicpony: Learning articulated 3d animals in the wild
Shangzhe Wu, Ruining Li, Tomas Jakab, Christian Rupprecht, and Andrea Vedaldi · 2022
Later among the works it cites.
Hi-lassie: High-fidelity articulated shape and skeleton discovery from sparse image ensemble
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang, and Varun Jampani · 2022
Later among the works it cites.
Lassie: Learning articulated shapes from sparse image ensemble via 3d part discovery
Chun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein, Ming-Hsuan Yang, and Varun Jampani · 2022
Later among the works it cites.
Farm3d: Learning articulated 3d animals by distilling 2d diffusion
Tomas Jakab, Ruining Li, Shangzhe Wu, Christian Rupprecht, and Andrea Vedaldi · 2023
Closest in time.
Dreambooth3d: Subject-driven text-to-3d generation
Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan Barron, et al · 2023
Closest in time.
Texture: Text-guided texturing of 3d shapes
Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or · 2023
Closest in time.
Text-to-4d dynamic scene generation
Uriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual, Iurii Makarov, Filippos Kokkinos, Naman Goyal, Andrea Vedaldi, Devi Parikh, Justin Johnson, et al · 2023
Closest in time.
Reconstructing animatable categories from videos
Gengshan Yang, Chaoyang Wang, N Dinesh Reddy, and Deva Ramanan · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang and Maneesh Agrawala · 2023
Closest in time.