Fetching the paper…
Reading the bibliography…
Answering questions about the spatial properties of the environment poses challenges for existing language and vision foundation models due to a lack of understanding of the 3D world notably in terms of relationships between objects.
2001
Earlier work this paper cites.
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Nießner, M.: Scannet: Richly-annotated 3d reconstructions of indoor scenes. In: Proc. Computer Vision and Pattern Recognition (CVPR), IEEE (2017)
2017
Earlier work this paper cites.
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., van den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3674–3683 (2018). https://doi.org/10.1109/CVPR.2018.00387
2018
Earlier work this paper cites.
Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuscenes: A multimodal dataset for autonomous driving. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 11618–11628 (2019), https://api.semanticscholar.org/CorpusID:85517967
2019
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 4171–4186. Association for Computational Linguistics, Minneapolis, Minnesota (Jun 2019). https://doi.org/10.18653/v1/N19-1423, https://aclanthology.org/N19-1423
2019
Earlier work this paper cites.
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D.: Habitat: A Platform for Embodied AI Research. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Wald, J., Avetisyan, A., Navab, N., Tombari, F., Niessner, M.: Rio: 3d object instance re-localization in changing indoor environments. In: Proceedings IEEE International Conference on Computer Vision (ICCV) (2019)
2019
Earlier work this paper cites.
Chen, D.Z., Chang, A.X., Nießner, M.: Scanrefer: 3d object localization in rgb-d scans using natural language. 16th European Conference on Computer Vision (ECCV) (2020)
2020
Earlier work this paper cites.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.t., Rocktäschel, T., Riedel, S., Kiela, D.: Retrieval-augmented generation for knowledge-intensive nlp tasks. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems. vol. 33, pp. 9459–9474. Curran Associates, Inc. (2020), https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf
2020
Earlier work this paper cites.
Mirzaee, R., Rajaby Faghihi, H., Ning, Q., Kordjamshidi, P.: SPARTQA: A textual question answering benchmark for spatial reasoning. In: Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., Zhou, Y. (eds.) Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 4582–4598. Association for Computational Linguistics, Online (Jun 2021). https://doi.org/10.18653/v1/2021.naacl-main.364, https://aclanthology.org/2021.naacl-main.364
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: Meila, M., Zhang, T. (eds.) Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 139, pp. 8748–8763. PMLR (18–24 Jul 2021), https://proceedings.mlr.press/v139/radford21a.html
2021
Earlier work this paper cites.
Azuma, D., Miyanishi, T., Kurita, S., Kawanabe, M.: Scanqa: 3d question answering for spatial scene understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Earlier work this paper cites.
Gu, J., Stefani, E., Wu, Q., Thomason, J., Wang, X.: Vision-and-language navigation: A survey of tasks, methods, and future directions. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics (2022). https://doi.org/10.18653/v1/2022.acl-long.524, http://dx.doi.org/10.18653/v1/2022.acl-long.524
2022
Earlier work this paper cites.
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: International Conference on Learning Representations (2022), https://openreview.net/forum?id=nZeVKeeFYf9
2022
Earlier work this paper cites.
Liu, J.: LlamaIndex (11 2022). https://doi.org/10.5281/zenodo.1234, https://github.com/jerryjliu/llama_index
2022
Earlier work this paper cites.
Ye, S., Chen, D., Han, S., Liao, J.: 3d question answering. IEEE Transactions on Visualization and Computer Graphics 30
2022
Cited alongside, same era.
2023
Cited alongside, same era.
Chang, H., Boyalakuntla, K., Lu, S., Cai, S., Jing, E.P., Keskar, S., Geng, S., Abbas, A., Zhou, L., Bekris, K., Boularious, A.: Context-aware entity grounding with open-vocabulary 3d scene graphs. In: 7th Annual Conference on Robot Learning (2023), https://openreview.net/forum?id=cjEI5qXoT0
2023
Cited alongside, same era.
Bozkir, E., Özdel, S., Lau, K.H.C., Wang, M., Gao, H., Kasneci, E.: Embedding large language models into extended reality: Opportunities and challenges for inclusion, engagement, and privacy. In: ACM Conversational User Interfaces 2024. CUI ’24, ACM (Jul 2024). https://doi.org/10.1145/3640794.3665563, http://dx.doi.org/10.1145/3640794.3665563
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Ma, X., Bhalgat, Y., Smart, B., Chen, S., Li, X., Ding, J., Gu, J., Chen, D., Peng, S., Bian, J., Torr, P., Pollefeys, M., Nießner, M., Reid, I., Chang, A., Laina, I., Prisacariu, V.: When llms step into the 3d world: a survey and meta-analysis of 3d tasks via multi-modal large language models. IEEE (2024)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gu, Q., Kuwajerwala, A., Morin, S., Jatavallabhula, K.M., Sen, B., Agarwal, A., Rivera, C., Paul, W., Ellis, K., Chellappa, R., Gan, C., de Melo, C.M., Tenenbaum, J.B., Torralba, A., Shkurti, F., Paull, L.: Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning (2023)
2023
Cited alongside, same era.
Hong, Y., Lin, C., Du, Y., Chen, Z., Tenenbaum, J.B., Gan, C.: 3d concept learning and reasoning from multi-view images. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2023)
2023
Cited alongside, same era.
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., Kasneci, G.: Chatgpt for good? on opportunities and challenges of large language models for education. Learning and Individual Differences 103
2023
Cited alongside, same era.
Li, M., Chen, X., Zhang, C., Chen, S., Zhu, H., Yin, F., Yu, G., Chen, T.: M3dbench: Let’s instruct large models with multi-modal 3d prompts (2023)
2023
Cited alongside, same era.
Ma, X., Yong, S., Zheng, Z., Li, Q., Liang, Y., Zhu, S.C., Huang, S.: Sqa3d: Situated question answering in 3d scenes. In: International Conference on Learning Representations (2023), https://openreview.net/forum?id=IDJx97BC38
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Qian, T., Chen, J., Zhuo, L., Jiao, Y., Jiang, Y.: Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario. In: AAAI Conference on Artificial Intelligence (2023), https://api.semanticscholar.org/CorpusID:258866014
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Yan, X., Yuan, Z., Du, Y., Liao, Y., Guo, Y., Cui, S., Li, Z.: Comprehensive visual question answering on point clouds through compositional scene manipulation. IEEE Transactions on Visualization & Computer Graphics (01), 1–13 (2023)
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
Microsoft: Semantic kernel (2024), https://github.com/microsoft/semantic-kernel , accessed: 2024-08-01
2024
Closest in time.
Miriam Schmidts, Nicholas M. Giner.: Understanding the basics: Introduction to the language of spatial analysis. https://proceedings.esri.com/library/userconf/proc18/tech-workshops/tw_1593-380.pdf , accessed: 2024-07-29
2024
Closest in time.
2024
Closest in time.
OpenAI: Gpt-4v(ision) system card (2023), https://api.semanticscholar.org/CorpusID:263218031 , accessed: 2024-08-01
2024
Closest in time.
OpenAI: New and improved embedding model (2024), https://openai.com/index/new-and-improved-embedding-model/ , accessed: 2024-08-01
2024
Closest in time.
Qi, Z., Fang, Y., Sun, Z., Wu, X., Wu, T., Wang, J., Lin, D., Zhao, H.: Gpt4point: A unified framework for point-language understanding and generation. In: CVPR (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Yuan, Z., Ren, J., Feng, C.M., Zhao, H., Cui, S., Li, Z.: Visual programming for zero-shot open-vocabulary 3d visual grounding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 20623–20633 (June 2024)
2024
Closest in time.