Fetching the paper…
Reading the bibliography…
This paper proposes a novel scene understanding task called Visual Jenga.
The Direction of Time
Hans Reichenbach · 1956
Earlier work this paper cites.
Machine perception of three-dimensional solids
Lawrence G Roberts · 1963
Earlier work this paper cites.
On the semantics of a glance at a scene
Irving Biederman · 1981
Earlier work this paper cites.
Causality: Models, reasoning, and inference
Judea Pearl · 2000
Earlier work this paper cites.
Recovering surface layout from an image
Derek Hoiem, Alexei A Efros, and Martial Hebert · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Beyond categories: The visual memex model for reasoning about object relationships
Tomasz Malisiewicz and Alyosha Efros · 2009
Earlier work this paper cites.
Blocks world revisited: Image understanding using qualitative geometry and mechanics
Abhinav Gupta, Alexei A. Efros, and Martial Hebert · 2010
Earlier work this paper cites.
Recovering occlusion boundaries from an image
Derek Hoiem, Alexei A. Efros, and Martial Hebert · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Support surface prediction in indoor scenes
Ruiqi Guo and Derek Hoiem · 2013
Earlier work this paper cites.
Scene collaging: Analysis and synthesis of natural images with semantic layers
Phillip Isola and Ce Liu · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael S Bernstein, and Fei-Fei Li · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Earlier work this paper cites.
Discovering causal signals in images
David Lopez-Paz, Robert Nishihara, Soumith Chintala, Bernhard Schölkopf, and Léon Bottou · 2016
Earlier work this paper cites.
Distinguishing cause from effect using observational data: methods and benchmarks
Joris M Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, and Bernhard Schölkopf · 2016
Earlier work this paper cites.
Visual stability prediction for robotic manipulation
Wenbin Li, Ales Leonardis, and Mario Fritz · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Earlier work this paper cites.
On support relations and semantic scene graphs
Michael Ying Yang, Wentong Liao, Hanno Ackermann, and Bodo Rosenhahn · 2017
Earlier work this paper cites.
Large-scale celebfaces attributes (celeba) dataset
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2018
Earlier work this paper cites.
Counterfactual image networks
Deniz Oktay, Carl Vondrick, and Antonio Torralba · 2018
Earlier work this paper cites.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation
Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
Object-driven multi-layer scene decomposition from a single image
Helisa Dhamo, Nassir Navab, and Federico Tombari · 2019
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Multi-object representation learning with iterative variational inference
Klaus Greff, Raphaël Lopez Kaufman, Rishabh Kabra, Nick Watters, Christopher Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, and Alexander Lerchner · 2019
Cited alongside, same era.
InstructPix2Pix: Learning to follow image editing instructions
T Brooks, A Holynski, and A A Efros · 2023
Later among the works it cites.
Generative models: What do they know? do they know things? let’s find out!
Xiaodan Du, Nicholas Kolkin, Greg Shakhnarovich, and Anand Bhattad · 2023
Later among the works it cites.
Diffusion self-guidance for controllable image generation
Dave Epstein, Allan Jabri, Ben Poole, Alexei Efros, and Aleksander Holynski · 2023
Later among the works it cites.
Your diffusion model is secretly a zero-shot classifier
Alexander C Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak · 2023
Later among the works it cites.
Differentiable blocks world: Qualitative 3D decomposition by rendering primitives
Tom Monnier, Jake Austin, Angjoo Kanazawa, Alexei A Efros, and Mathieu Aubry · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Panoptic feature pyramid networks
Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar · 2019
Cited alongside, same era.
How to make a pizza: Learning a compositional layer-based gan model
Dim P Papadopoulos, Youssef Tamaazousti, Ferda Ofli, Ingmar Weber, and Antonio Torralba · 2019
Cited alongside, same era.
Counterfactuals uncover the modular structure of deep generative models
Michel Besserve, Arash Mehrjou, Rémy Sun, and Bernhard Schölkopf · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
ACRONYM: A large-scale grasp dataset based on simulation
Clemens Eppner, Arsalan Mousavian, and Dieter Fox · 2021
Cited alongside, same era.
Probabilistic Causation
Christopher Hitchcock · 2021
Cited alongside, same era.
Variational diffusion models
Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Cited alongside, same era.
Vaibhav Vavilala, Seemandhar Jain, Rahul Vasanth, Anand Bhattad, and David Forsyth · 2023
Later among the works it cites.
Mental jenga: A counterfactual simulation model of causal judgments about physical support
Liang Zhou, Kevin A Smith, Joshua B Tenenbaum, and Tobias Gerstenberg · 2023
Later among the works it cites.
Stylitgan: Image-based relighting via latent control
Anand Bhattad, James Soole, and David A Forsyth · 2024
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Later among the works it cites.
EraseDraw: Learning to insert objects by erasing them from images
Alper Canberk, Maksym Bondarenko, Ege Ozguroglu, Ruoshi Liu, and Carl Vondrick · 2024
Later among the works it cites.
Physically grounded vision-language models for robotic manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh · 2024
Later among the works it cites.
Visual hallucinations of multi-modal large language models
Wen Huang, Hongbin Liu, Minxin Guo, and Neil Zhenqiang Gong · 2024
Later among the works it cites.
Repurposing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler · 2024
Later among the works it cites.
Object-level scene deocclusion
Zhengzhe Liu, Qing Liu, Chirui Chang, Jianming Zhang, Daniil Pakhomov, Haitian Zheng, Zhe Lin, Daniel Cohen-Or, and Chi-Wing Fu · 2024
Later among the works it cites.
Object remover performance evaluation methods using class-wise object removal images
Changsuk Oh, Dongseok Shim, Taekbeom Lee, and H Jin Kim · 2024
Later among the works it cites.
DINOv2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Later among the works it cites.
Pix2gestalt: Amodal segmentation by synthesizing wholes
Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, and Carl Vondrick · 2024
Later among the works it cites.
ObjectDrop: Bootstrapping counterfactuals for photorealistic object removal and insertion
Daniel Winter, Matan Cohen, Shlomi Fruchter, Yael Pritch, Alex Rav-Acha, and Yedid Hoshen · 2024
Later among the works it cites.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
Later among the works it cites.
A task is worth one word: Learning with task prompts for high-quality versatile image inpainting
Junhao Zhuang, Yanhong Zeng, Wenran Liu, Chun Yuan, and Kai Chen · 2024
Later among the works it cites.
Adobe firefly - free generative AI for creatives
Adobe Inc · 2025
Closest in time.
SAM 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer · 2025
Closest in time.
A general protocol to probe large vision models for 3D physical understanding
Guanqi Zhan, Chuanxia Zheng, Weidi Xie, and Andrew Zisserman · 2025
Closest in time.