Fetching the paper…
Reading the bibliography…
A critical aspect of human visual perception is the ability to parse visual scenes into individual objects and further into object parts, forming part-whole hierarchies.
Principles of gestalt psychology
O. Reiser and K. Koffka · 1936
Earlier work this paper cites.
Modeling by example
Thomas Funkhouser, Michael Kazhdan, Philip Shilane, Patrick Min, William Kiefer, Ayellet Tal, Szymon Rusinkiewicz, and David Dobkin · 2004
Earlier work this paper cites.
Manipulating articulated objects with interactive perception
Dov Katz and Oliver Brock · 2008
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan · 2009
Earlier work this paper cites.
Consistent segmentation of 3d models
Aleksey Golovinskiy and Thomas Funkhouser · 2009
Earlier work this paper cites.
Learning 3d mesh segmentation and labeling
Evangelos Kalogerakis, Aaron Hertzmann, and Karan Singh · 2010
Earlier work this paper cites.
Probabilistic reasoning for assembly-based 3d modeling
Siddhartha Chaudhuri, Evangelos Kalogerakis, Leonidas Guibas, and Vladlen Koltun · 2011
Earlier work this paper cites.
Joint shape segmentation with linear programming
Qixing Huang, Vladlen Koltun, and Leonidas Guibas · 2011
Earlier work this paper cites.
A probabilistic model for component-based shape synthesis
Evangelos Kalogerakis, Siddhartha Chaudhuri, Daphne Koller, and Vladlen Koltun · 2012
Earlier work this paper cites.
Learning part-based templates from large collections of 3d shapes
Vladimir G Kim, Wilmot Li, Niloy J Mitra, Siddhartha Chaudhuri, Stephen DiVerdi, and Thomas Funkhouser · 2013
Earlier work this paper cites.
Meta-representation of shape families
Noa Fish, Melinos Averkiou, Oliver Van Kaick, Olga Sorkine-Hornung, Daniel Cohen-Or, and Niloy J Mitra · 2014
Earlier work this paper cites.
Diagram understanding in geometry questions
Minjoon Seo, Hannaneh Hajishirzi, Ali Farhadi, and Oren Etzioni · 2014
Earlier work this paper cites.
3d shape segmentation and labeling via extreme learning machine
Zhige Xie, Kai Xu, Ligang Liu, and Yueshan Xiong · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Semantic shape editing using deformation handles
Mehmet Ersin Yumer, Siddhartha Chaudhuri, Jessica K Hodgins, and Levent Burak Kara · 2015
Earlier work this paper cites.
Learning how objects function via co-analysis of interactions
Ruizhen Hu, Oliver van Kaick, Bojian Wu, Hui Huang, Ariel Shamir, and Hao Zhang · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Yuke Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, Stephanie Chen, Yannis Kalantidis, L. Li, D. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
A scalable active framework for region annotation in 3d shape collections
Li Yi, Vladimir G Kim, Duygu Ceylan, I-Chao Shen, Mengyan Yan, Hao Su, Cewu Lu, Qixing Huang, Alla Sheffer, and Leonidas Guibas · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, O. Groth, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
VQS: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
Chuang Gan, Yandong Li, Haoxiang Li, Chen Sun, and Boqing Gong · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick · 2017
Cited alongside, same era.
Learning to predict part mobility from a single static snapshot
Ruizhen Hu, Wenchao Li, Oliver Van Kaick, Ariel Shamir, Hao Zhang, and Hui Huang · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, Bharath Hariharan, L. V. D. Maaten, Li Fei-Fei, C. L. Zitnick, and Ross B. Girshick · 2017
Cited alongside, same era.
Shape2motion: Joint analysis of motion parts and attributes from 3d shapes
Xiaogang Wang, Bin Zhou, Yahao Shi, Xiaowu Chen, Qinping Zhao, and Kai Xu · 2019
Later among the works it cites.
Sagnet: Structure-aware generative network for 3d-shape modeling
Zhijie Wu, Xiang Wang, Di Lin, Dani Lischinski, Daniel Cohen-Or, and Hui Huang · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Later among the works it cites.
Raven: A dataset for relational and analogical visual reasoning
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu · 2019
Later among the works it cites.
Compositionally generalizable 3d structure prediction
Songfang Han, Jiayuan Gu, Kaichun Mo, Li Yi, Siyu Hu, Xuejin Chen, and Hao Su · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evangelos Kalogerakis, Melinos Averkiou, Subhransu Maji, and Siddhartha Chaudhuri · 2017
Cited alongside, same era.
Grass: Generative recursive autoencoders for shape structures
Jun Li, Kai Xu, Siddhartha Chaudhuri, Ersin Yumer, Hao Zhang, and Leonidas Guibas · 2017
Cited alongside, same era.
Complementme: Weakly-supervised component suggestions for 3d modeling
Minhyuk Sung, Hao Su, Vladimir G Kim, Siddhartha Chaudhuri, and Leonidas Guibas · 2017
Cited alongside, same era.
Learning hierarchical shape segmentation and labeling from online repositories
Li Yi, Leonidas Guibas, Aaron Hertzmann, Vladimir G Kim, Hao Su, and Ersin Yumer · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Linking wordnet to 3d shapes
Angel X Chang, Rishi Mago, Pranav Krishna, Manolis Savva, and Christiane Fellbaum · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
D. A. Hudson and Christopher D. Manning · 2018
Cited alongside, same era.
Closed loop neural-symbolic learning via integrating neural perception, grammar parsing, and symbolic reasoning
Qing Li, Siyuan Huang, Yining Hong, Yixin Chen, Ying Nian Wu, and Song-Chun. Zhu · 2020
Later among the works it cites.
Object-centric learning with slot attention, 2020
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Learning to group: a bottom-up framework for 3d part discovery in unseen categories
Tiange Luo, Kaichun Mo, Zhiao Huang, Jiarui Xu, Siyu Hu, Liwei Wang, and Hao Su · 2020
Later among the works it cites.
Structedit: Learning structural shape variations
Kaichun Mo, Paul Guerrero, Li Yi, Hao Su, Peter Wonka, Niloy J Mitra, and Leonidas J Guibas · 2020
Later among the works it cites.
Sapien: A simulated part-based interactive environment
Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, et al · 2020
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, P. Kohli, Jiajun Wu, A. Torralba, and J. Tenenbaum · 2020
Later among the works it cites.
The gap of semantic parsing: A survey on automatic math word problem solvers
D. Zhang, Lei Wang, Nuo Xu, B. Dai, and H. Shen · 2020
Later among the works it cites.
Grounding physical concepts of objects and events through dynamic visual reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B Tenenbaum, and Chuang Gan · 2021
Closest in time.
Generative scene graph networks
Fei Deng, Zhuo Zhi, Donghun Lee, and Sungjin Ahn · 2021
Closest in time.
Threedworld: A platform for interactive multi-modal physical simulation
Chuang Gan, Jeremy Schwartz, Seth Alter, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, et al · 2021
Closest in time.
How to represent part-whole hierarchies in a neural network
Geoffrey Hinton · 2021
Closest in time.
Mdetr - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Y. LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Closest in time.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu · 2021
Closest in time.
Language-mediated, object-centric representation learning
Ruocheng Wang, Jiayuan Mao, Samuel J. Gershman, and Jiajun Wu · 2021
Closest in time.