Fetching the paper…
Reading the bibliography…
Inverse graphics -- the task of inverting an image into physical variables that, when rendered, enable reproduction of the observed scene -- is a fundamental challenge in computer vision and graphics.
Machine perception of three-dimensional solids
Lawrence G. Roberts · 1963
Earlier work this paper cites.
Geometric modeling for computer vision
Bruce G. Baumgart · 1974
Earlier work this paper cites.
Visual perception by computer
Thomas Binford · 1975
Earlier work this paper cites.
Lectures in Pattern Theory I, II and III: Pattern Analysis, Pattern Synthesis and Regular Structures
Ulf Grenander · 1981
Earlier work this paper cites.
Introduction: A Bayesian formulation of visual perception
David C. Knill, Daniel Kersten, and Alan Yuille · 1996
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Vision as Bayesian inference: Analysis by synthesis?
Alan Yuille and Daniel Kersten · 2006
Earlier work this paper cites.
Recovering the spatial layout of cluttered rooms
Varsha Hedau, Derek Hoiem, and David Forsyth · 2009
Earlier work this paper cites.
Geometric reasoning for single image structure recovery
David C. Lee, Martial Hebert, and Takeo Kanade · 2009
Earlier work this paper cites.
EPnP: An accurate O(n) solution to the PnP problem
Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua · 2009
Earlier work this paper cites.
A connection between partial symmetry and inverse procedural modeling
Martin Bokeloh, Michael Wand, and Hans-Peter Seidel · 2010
Earlier work this paper cites.
Inverse procedural modeling by automatic generation of L-systems
Ondrej Št’ava, Bedrich Beneš, Radomir Měch, Daniel G. Aliaga, and Peter Krištof · 2010
Earlier work this paper cites.
Approximate Bayesian image interpretation using generative probabilistic graphics programs
Vikash K. Mansinghka, Tejas D. Kulkarni, Yura N. Perov, and Josh Tenenbaum · 2013
Earlier work this paper cites.
SLAM++: Simultaneous localisation and mapping at the level of objects
Renato F. Salas-Moreno, Richard A. Newcombe, Hauke Strasdat, Paul HJ Kelly, and Andrew J. Davison · 2013
Earlier work this paper cites.
Seeing 3D chairs: Exemplar part-based 2D-3D alignment using a large dataset of CAD models
Mathieu Aubry, Daniel Maturana, Alexei A. Efros, Bryan C. Russell, and Josef Sivic · 2014
Earlier work this paper cites.
FPM: Fine pose parts-based model with 3D CAD models
Joseph J. Lim, Aditya Khosla, and Antonio Torralba · 2014
Earlier work this paper cites.
Latent-class Hough forests for 3D object detection and pose estimation
Alykhan Tejani, Danhang Tang, Rigas Kouskouridas, and Tae-Kyun Kim · 2014
Earlier work this paper cites.
Inverse procedural modelling of trees
Ondrej Št’ava, Soren Pirk, Julian Kratt, Baoquan Chen, Radomir Měch, Oliver Deussen, and Bedrich Beneš · 2014
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
ShapeNet: An information-rich 3D model repository, 2015
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Picture: A probabilistic programming language for scene perception
Tejas D. Kulkarni, Pushmeet Kohli, Joshua B. Tenenbaum, and Vikash Mansinghka · 2015
Earlier work this paper cites.
Learning informative edge maps for indoor scene layout prediction
Arun Mallya and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Viewpoints and keypoints
Shubham Tulsiani and Jitendra Malik · 2015
Earlier work this paper cites.
Part-based modelling of compound scenes from images
Anton van den Hengel, Chris Russell, Anthony Dick, John Bastian, Daniel Pooley, Lachlan Fleming, and Lourdes Agapito · 2015
Earlier work this paper cites.
Marr revisited: 2D-3D alignment via surface normal prediction
Aayush Bansal, Bryan C. Russell, and Abhinav Gupta · 2016
Earlier work this paper cites.
3D-R2N2: A unified approach for single and multi-view 3D object reconstruction
Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese · 2016
Earlier work this paper cites.
DeLay: Robust spatial layout estimation for cluttered indoor scenes
Saumitro Dasgupta, Kuan Fang, Kevin Chen, and Silvio Savarese · 2016
Earlier work this paper cites.
A point set generation network for 3D object reconstruction from a single image
Haoqiang Fan, Hao Su, and Leonidas J. Guibas · 2017
Earlier work this paper cites.
Program synthesis
Sumit Gulwani, Oleksandr Polozov, and Rishabh Singh · 2017
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
6-DOF object pose from semantic keypoints
Georgios Pavlakos, Xiaowei Zhou, Aaron Chan, Konstantinos G. Derpanis, and Kostas Daniilidis · 2017
Earlier work this paper cites.
A coarse-to-fine indoor layout estimation (CFILE) method
Yuzhuo Ren, Shangwen Li, Chen Chen, and C-C Jay Kuo · 2017
Earlier work this paper cites.
Learning shape abstractions by assembling volumetric primitives
Shubham Tulsiani, Hao Su, Leonidas J. Guibas, Alexei A. Efros, and Jitendra Malik · 2017
Earlier work this paper cites.
Neural scene de-rendering
Jiajun Wu, Joshua B. Tenenbaum, and Pushmeet Kohli · 2017
Earlier work this paper cites.
Blender – A 3D modelling and rendering package
Blender · 2018
Earlier work this paper cites.
InverseCSG: Automatic conversion of 3D models to CSG trees
Tao Du, Jeevana Priya Inala, Yewen Pu, Andrew Spielberg, Adriana Schulz, Daniela Rus, Armando Solar-Lezama, and Wojciech Matusik · 2018
Earlier work this paper cites.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Earlier work this paper cites.
Synthesizing programs for images using reinforced adversarial learning
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, S.M. Ali Eslami, and Oriol Vinyals · 2018
Earlier work this paper cites.
A papier-mâché approach to learning 3D surface generation
Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan C. Russell, and Mathieu Aubry · 2018
Earlier work this paper cites.
Holistic 3D scene parsing and reconstruction from a single RGB image
Siyuan Huang, Siyuan Qi, Yixin Zhu, Yinxue Xiao, Yuanlu Xu, and Song-Chun Zhu · 2018
Earlier work this paper cites.
3D-RCNN: Instance-level 3D object reconstruction via render-and-compare
Abhijit Kundu, Yin Li, and James M. Rehg · 2018
Cited alongside, same era.
Pixel2Mesh: Generating 3D mesh models from single RGB images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang · 2018
Cited alongside, same era.
PoseCNN: A convolutional neural network for 6D object pose estimation in cluttered scenes, 2018
Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox · 2018
Cited alongside, same era.
3D-aware scene manipulation via inverse graphics
Shunyu Yao, Tzu Ming Hsu, Jun-Yan Zhu, Jiajun Wu, Antonio Torralba, Bill Freeman, and Josh Tenenbaum · 2018
Cited alongside, same era.
Neural-symbolic VQA: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Cited alongside, same era.
Write, execute, assess: Program synthesis with a REPL
Inferring CAD modeling sequences using zone graphs
Xianghao Xu, Wenzhe Peng, Chin-Yi Cheng, Karl D.D. Willis, and Daniel Ritchie · 2021
Later among the works it cites.
Holistic 3D scene understanding from a single image with implicit representation
Cheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng, Marc Pollefeys, and Shuaicheng Liu · 2021
Later among the works it cites.
Flamingo: A visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Unsupervised learning of shape programs with repeatable implicit parts
Boyang Deng, Sumith Kulal, Zhengyang Dong, Congyue Deng, Yonglong Tian, and Jiajun Wu · 2022
Later among the works it cites.
Learning 3D object shape and layout without 3D supervision
Georgia Gkioxari, Nikhila Ravi, and Justin Johnson · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Ellis, Maxwell Nye, Yewen Pu, Felix Sosa, Josh Tenenbaum, and Armando Solar-Lezama · 2019
Cited alongside, same era.
Mesh R-CNN
Georgia Gkioxari, Jitendra Malik, and Justin Johnson · 2019
Cited alongside, same era.
Meta-Sim: Learning to generate synthetic datasets
Amlan Kar, Aayush Prakash, Ming-Yu Liu, Eric Cameracci, Justin Yuan, Matt Rusiniak, David Acuna, Antonio Torralba, and Sanja Fidler · 2019
Cited alongside, same era.
Learning to describe scenes with programs
Yunchao Liu, Jiajun Wu, Zheng Wu, Daniel Ritchie, William T. Freeman, and Joshua B. Tenenbaum · 2019
Cited alongside, same era.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu · 2019
Cited alongside, same era.
Occupancy networks: Learning 3D reconstruction in function space
Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger · 2019
Cited alongside, same era.
DeepSDF: Learning continuous signed distance functions for shape representation
Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove · 2019
Cited alongside, same era.
Can Gümeli, Angela Dai, and Matthias Nießner · 2022
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Later among the works it cites.
PLAD: Learning to infer shape programs with pseudo-labels and approximate distributions
R. Kenny Jones, Homer Walke, and Daniel Ritchie · 2022
Later among the works it cites.
Panoptic neural fields: A semantic object-aware neural scene representation
Abhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi, Caroline Pantofaru, Leonidas J. Guibas, Andrea Tagliasacchi, Frank Dellaert, and Thomas Funkhouser · 2022
Later among the works it cites.
Towards high-fidelity single-view holistic reconstruction of indoor scenes
Haolin Liu, Yujian Zheng, Guanying Chen, Shuguang Cui, and Xiaoguang Han · 2022
Later among the works it cites.
Robust category-level 6D pose estimation with coarse-to-fine rendering of neural features
Wufei Ma, Angtian Wang, Alan Yuille, and Adam Kortylewski · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
ExtrudeNet: Unsupervised inverse sketch-and-extrude for shape parsing
Daxuan Ren, Jianmin Zheng, Jianfei Cai, Jiatong Li, and Junzhe Zhang · 2022
Later among the works it cites.
pOp: Parameter optimization of differentiable vector patterns
Marzia Riso, Davide Sforza, and Fabio Pellacini · 2022
Later among the works it cites.
Vitruvion: A generative model of parametric CAD sketches
Ari Seff, Wenda Zhou, Nick Richardson, and Ryan P. Adams · 2022
Later among the works it cites.
CAPRI-Net: Learning compact CAD shapes with adaptive primitive assembly
Fenggen Yu, Zhiqin Chen, Manyi Li, Aditya Sanghi, Hooman Shayani, Ali Mahdavi-Amiri, and Hao Zhang · 2022
Later among the works it cites.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, et al · 2023
Later among the works it cites.
Generating context-aware natural answers for questions in 3D scenes
Mohammed Munzer Dwedari, Matthias Niessner, and Zhenyu Chen · 2023
Later among the works it cites.
Improving unsupervised visual program inference with code rewriting families
Aditya Ganeshan, R. Kenny Jones, and Daniel Ritchie · 2023
Later among the works it cites.
3d-llm: Injecting the 3d world into large language models
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan · 2023
Later among the works it cites.
ShapeCoder: Discovering abstractions for visual programs from unstructured primitives
R. Kenny Jones, Paul Guerrero, Niloy J. Mitra, and Daniel Ritchie · 2023
Later among the works it cites.
ReparamCAD: Zero-shot CAD re-parameterization for interactive manipulation
Milin Kodnongbua, Benjamin Jones, Maaz Bin Safeer Ahmad, Vladimir Kim, and Adriana Schulz · 2023
Later among the works it cites.
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov · 2023
Later among the works it cites.
Super-CLEVR: A virtual benchmark to diagnose domain robustness in visual reasoning
Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, and Alan L. Yuille · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Differentiable blocks world: Qualitative 3D decomposition by rendering primitives
Tom Monnier, Jake Austin, Angjoo Kanazawa, Alexei A. Efros, and Mathieu Aubry · 2023
Later among the works it cites.
Learning 3D scene priors with 2D supervision
Yinyu Nie, Angela Dai, Xiaoguang Han, and Matthias Nießner · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Instruction tuning with GPT-4, 2023
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Later among the works it cites.
Neurosymbolic models for computer graphics
Daniel Ritchie, Paul Guerrero, R. Kenny Jones, Niloy J. Mitra, Adriana Schulz, Karl D. D. Willis, and Jiajun Wu · 2023
Later among the works it cites.
ProgPrompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2023
Later among the works it cites.
3D-GPT: Procedural 3D modeling with large language models
Chunyi Sun, Junlin Han, Weijian Deng, Xinlong Wang, Zishan Qin, and Stephen Gould · 2023
Later among the works it cites.
Stanford Alpaca: An instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Convex decomposition of indoor scenes
Vaibhav Vavilala and David Forsyth · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
ULIP: Learning a unified representation of language, images, and point clouds for 3D understanding
Le Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín, Jiajun Wu, Caiming Xiong, Ran Xu, Juan Carlos Niebles, and Silvio Savarese · 2023
Later among the works it cites.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
YOLO-6D-Pose: Enhancing YOLO for single-stage monocular multi-object 6D pose estimation
Debapriya Maji, Soyeb Nagori, Manu Mathew, and Deepak Poddar · 2024
Closest in time.