Fetching the paper…
Reading the bibliography…
We study open-world 3D scene understanding, a family of tasks that require agents to reason about their 3D environment with an open-set vocabulary and out-of-domain visual inputs - a critical skill for robots to operate in the unstructured 3D world.
Visual perception by computer
I. Binford · 1971
Earlier work this paper cites.
“what” and “where” in spatial language and spatial cognition
L. Barbara and R. Jackendoff · 1993
Earlier work this paper cites.
Spatial language and spatial representation
W. G. Hayward and M. J. Tarr · 1995
Earlier work this paper cites.
Blocks world revisited: Image understanding using qualitative geometry and mechanics
A. Gupta, A. A. Efros, and M. Hebert · 2010
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
Sliding Shapes for 3D object detection in depth images
Shuran Song and J. Xiao · 2014
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
S. Song, S. P. Lichtenberg, and J. Xiao · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al · 2015
Earlier work this paper cites.
Predicting complete 3d models of indoor scenes
R. Guo, C. Zou, and D. Hoiem · 2015
Earlier work this paper cites.
Deep sliding shapes for amodal 3D object detection in rgb-d images
Shuran Song and J. Xiao · 2016
Earlier work this paper cites.
3d u-net: Learning dense volumetric segmentation from sparse annotation
Ö. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger · 2016
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image
S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser · 2017
Earlier work this paper cites.
Learning shape abstractions by assembling volumetric primitives
S. Tulsiani, H. Su, L. J. Guibas, A. A. Efros, and J. Malik · 2017
Earlier work this paper cites.
Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans
A. Dai, D. Ritchie, M. Bokeloh, S. Reed, J. Sturm, and M. Nießner · 2018
Earlier work this paper cites.
3d-sis: 3d semantic instance segmentation of rgb-d scans
J. Hou, A. Dai, and M. Nießner · 2019
Earlier work this paper cites.
Scan2cad: Learning cad model alignment in rgb-d scans
A. Avetisyan, M. Dahnert, A. Dai, M. Savva, A. X. Chang, and M. Nießner · 2019
Earlier work this paper cites.
Cascaded context pyramid for full-resolution 3d semantic scene completion
P. Zhang, W. Liu, Y. Lei, H. Lu, and X. Yang · 2019
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Cited alongside, same era.
Habitat 2.0: Training home assistants to rearrange their habitat
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. Turner, N. Maestre, M. Mukadam, D. Chaplot, O. Maksymets, A. Gokaslan, V. Vondrus, S. Dharur, F. Meier, W. Galuba, A. Chang, Z. Kira, V. Koltun, J. Malik, M. Savva, and D. Batra · 2021
Transformer interpretability beyond attention visualization
H. Chefer, S. Gur, and L. Wolf · 2021
Later among the works it cites.
How much can clip benefit vision-and-language tasks?
S. Shen, L. H. Li, H. Tan, M. Bansal, A. Rohrbach, K.-W. Chang, Z. Yao, and K. Keutzer · 2021
Later among the works it cites.
Learning to compose visual relations
N. Liu, S. Li, Y. Du, J. Tenenbaum, and A. Torralba · 2021
Later among the works it cites.
Semantic scene completion via integrating instances and scene in-the-loop
Y. Cai, X. Chen, C. Zhang, K.-Y. Lin, X. Wang, and H. Li · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi · 2021
Cited alongside, same era.
Combined scaling for zero-shot transfer learning
H. Pham, Z. Dai, G. Ghiasi, H. Liu, A. W. Yu, M.-T. Luong, M. Tan, and Q. V. Le · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Robust fine-tuning of zero-shot models
M. Wortsman, G. Ilharco, M. Li, J. W. Kim, H. Hajishirzi, A. Farhadi, H. Namkoong, and L. Schmidt · 2021
Cited alongside, same era.
The evolution of out-of-distribution robustness throughout fine-tuning
A. Andreassen, Y. Bahri, B. Neyshabur, and R. Roelofs · 2021
Cited alongside, same era.
Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers
H. Chefer, S. Gur, and L. Wolf · 2021
Cited alongside, same era.
S. Subramanian, W. Merrill, T. Darrell, M. Gardner, S. Singh, and A. Rohrbach · 2022
Closest in time.
Clip on wheels: Zero-shot object navigation as object localization and exploration
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Denseclip: Language-guided dense prediction with context-aware prompting
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu · 2022
Closest in time.
Adapting clip for phrase localization without further training
J. Li, G. Shakhnarovich, and R. A. Yeh · 2022
Closest in time.
Clip models are few-shot learners: Empirical studies on vqa and visual entailment
H. Song, L. Dong, W.-N. Zhang, T. Liu, and F. Wei · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Closest in time.
Languagerefer: Spatial-language model for 3d visual grounding
J. Roh, K. Desingh, A. Farhadi, and D. Fox · 2022
Closest in time.
Optimizing relevance maps of vision transformers improves robustness
H. Chefer, I. Schwartz, and L. Wolf · 2022
Closest in time.
Rethinking attention-model explainability through faithfulness violation test
Y. Liu, H. Li, Y. Guo, C. Kong, J. Li, and S. Wang · 2022
Closest in time.
When and why vision-language models behave like bags-of-words, and what to do about it?, 2022
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou · 2022
Closest in time.
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt · 2022
Closest in time.