Fetching the paper…
Reading the bibliography…
Modern tools for class-agnostic image segmentation (e.g., SegmentAnything) and open-set semantic understanding (e.g., CLIP) provide unprecedented opportunities for robot perception and mapping.
A. M. Treisman and G. Gelade, “A feature-integration theory of attention,” in
1980
Earlier work this paper cites.
N. Slonim and N. Tishby, “Agglomerative information bottleneck,” in
1999
Earlier work this paper cites.
N. Tishby, F. Pereira, and W. Bialek, “The information bottleneck method,”
2001
Earlier work this paper cites.
S. Gordon, H. Greenspan, and J. Goldberger, “Applying the information bottleneck principle to unsupervised clustering of discrete and continuous image representations,” in
2003
Earlier work this paper cites.
R. Roberts, D.-N. Ta, J. Straub, and F. Dellaert, “Saliency detection and model-based tracking: a two part vision system for small robot navigation in forested environment,” in
2012
Earlier work this paper cites.
S. Soatto and A. Chiuso, “Visual scene representations: sufficiency, minimality, invariance and deep approximation,” in
2014
Earlier work this paper cites.
S. Soatto and A. Chiuso, “Visual representations: Defining properties and deep approximations,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
I. Armeni, Z. He, J. Gwak, A. Zamir, M. Fischer, J. Malik, and S. Savarese, “3D scene graph: A structure for unified semantics, 3D space, and camera,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
——, “Q-Tree Search: An Information-Theoretic Approach Toward Hierarchical Abstractions for Agents With Computational Limitations,”
2020
Earlier work this paper cites.
S. Wu, J. Wald, K. Tateno, N. Navab, and F. Tombari, “SceneGraphFusion: Incremental 3D scene graph prediction from RGB-D sequences,” in
2021
Earlier work this paper cites.
A. Radford
2021
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”
2021
Earlier work this paper cites.
D. T. Larsson, D. Maity, and P. Tsiotras, “Information-Theoretic Abstractions for Planning in Agents With Computational Constraints,”
2021
Earlier work this paper cites.
G. Ilharco
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Minderer
2022
Earlier work this paper cites.
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in
2022
Earlier work this paper cites.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in
2022
Cited alongside, same era.
H. Ha and S. Song, “Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models,” in
2022
Cited alongside, same era.
K. Jatavallabhula
2023
Cited alongside, same era.
A. Kirillov
2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in
2023
Cited alongside, same era.
OpenAI, “GPT-4 technical report,”
2023
Cited alongside, same era.
F. Taioli, F. Cunico, F. Girella, R. Bologna, A. Farinelli, and M. Cristani, “Language-enhanced rnr-map: Querying renderable neural radiance field maps with natural language,” in
2023
Later among the works it cites.
S. Peng, K. Genova, C. M. Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocabularies,” in
2023
Later among the works it cites.
J. Wang, J. J. Tarrio, L. de Agapito, P. F. Alcantarilla, and A. Vakhitov, “Semlaps: Real-time semantic mapping with latent prior networks and quasi-planar segmentation,”
2023
Later among the works it cites.
H. Chang
2023
Later among the works it cites.
A. Takmaz, E. Fedele, R. W. Sumner, M. Pollefeys, F. Tombari, and F. Engelmann, “OpenMask3D: Open-Vocabulary 3D Instance Segmentation,” in
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan, “3d-llm: Injecting the 3d world into large language models,”
2023
Cited alongside, same era.
C. Zhao, Y. Shen, Z. Chen, M. Ding, and C. Gan, “Textpsg: Panoptic scene graph generation from textual descriptions,” in
2023
Cited alongside, same era.
S. Garg, “Robohop: Segment-based topological map representation for open-world visual navigation,” in
2023
Cited alongside, same era.
C. Huang, O. Mees, A. Zeng, and W. Burgard, “Visual language maps for robot navigation,” in
2023
Cited alongside, same era.
Y. Wang, T. G. Rudner, and A. G. Wilson, “Visual explanations of image-text representations via multi-modal information bottleneck attribution,”
2023
Later among the works it cites.
C. Parameshwara
2023
Later among the works it cites.
L. Mur-Labadia, R. Martinez-Cantin, and J. J. Guerrero, “Bayesian deep learning for affordance segmentation in images,”
2023
Later among the works it cites.
L. Mur-Labadia, J. J. Guerrero, and R. Martinez-Cantin, “Multi-label affordance mapping from egocentric vision,” in
2023
Later among the works it cites.
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled feature fields enable few-shot language-guided manipulation,” in
2023
Later among the works it cites.
2024
Closest in time.
P. Sharma
2024
Closest in time.
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” in
2024
Closest in time.
C. M. Kim, M. Wu, J. Kerr, K. Goldberg, M. Tancik, and A. Kanazawa, “Garfield: Group anything with radiance fields,” in
2024
Closest in time.
S. Koch, P. Hermosilla, N. Vaskevicius, M. Colosi, and T. Ropinski, “Lang3dsg: Language-based contrastive pre-training for 3d scene graph prediction,” in
2024
Closest in time.
K. Yamazaki
2024
Closest in time.
C. Kassab, M. Mattamala, L. Zhang, and M. Fallon, “Language-extended indoor slam (lexis): A versatile system for real-time visual scene understanding,”
2024
Closest in time.
A. Werby, C. Huang, M. Büchner, A. Valada, and W. Burgard, “Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,”
2024
Closest in time.
A. Eftekhar, K.-H. Zeng, J. Duan, A. Farhadi, A. Kembhavi, and R. Krishna, “Selective visual representations improve convergence and generalization for embodied AI,” in
2024
Closest in time.