Fetching the paper…
Reading the bibliography…
We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
3d scene graph: A structure for unified semantics, 3d space, and camera
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese · 1910
Earlier work this paper cites.
3d dynamic scene graphs: Actionable spatial perception with places, objects, and humans
A. Rosinol, A. Gupta, M. Abate, J. Shi, and L. Carlone · 2002
Earlier work this paper cites.
Learning 3d semantic scene graphs from 3d indoor reconstructions
J. Wald, H. Dhamo, N. Navab, and F. Tombari · 2004
Earlier work this paper cites.
A benchmark for rgb-d visual odometry, 3d reconstruction and slam
A. Handa, T. Whelan, J. McDonald, and A. J. Davison · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
3-d scene graph: A sparse and semantic representation of physical environments for intelligent agents
U. H. Kim, J. M. Park, T. J. Song, and J. H. Kim · 2019
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam
C. Campos, R. Elvira, J. J. G. Rodríguez, J. M. Montiel, and J. D. Tardós · 2021
Cited alongside, same era.
Lit: Zero-shot transfer with locked-image text tuning
Detecting twenty-thousand classes using image-level supervision
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Later among the works it cites.
Simple open-vocabulary object detection with vision transformers. arxiv 2022
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, et al · 2022
Later among the works it cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2022
Later among the works it cites.
Findit: Generalized localization with natural language queries
W. Kuo, F. Bertsch, W. Li, A. Piergiovanni, M. Saffar, and A. Angelova · 2022
Later among the works it cites.
Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data
S. Lu, H. Chang, E. P. Jing, A. Boularias, and K. Bekris · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2022
Cited alongside, same era.
Language-driven semantic segmentation
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl · 2022
Cited alongside, same era.
Scaling open-vocabulary image segmentation with image-level labels
G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin · 2022
Cited alongside, same era.
Groupvit: Semantic segmentation emerges from text supervision
J. Xu, S. De Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, and X. Wang · 2022
Cited alongside, same era.
Characterizing structural relationships in scenes using graph kernels
M. Fisher, M. Savva, and P. Hanrahan
Cited in the paper.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui
Cited in the paper.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song · 2023
Closest in time.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3d mapping
K. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, J. Tenenbaum, C. de Melo, M. Krishna, L. Paull, F. Shkurti, and A. Torralba · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3d mapping
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, et al · 2023
Closest in time.