Fetching the paper…
Reading the bibliography…
3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments.
Wordnet: a lexical database for english
Miller, G.A., 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries, in: Text summarization branches out, pp. 74–81
Lin, C.Y., 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72
Banerjee, S., Lavie, A., 2005 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.B., 2020 · 2005
Earlier work this paper cites.
Models for multiparty engagement in open-world dialog, in: Proceedings of the SIGDIAL 2009 conference, the 10th annual meeting of the special interest group on discourse and dialogue, p. 10
Bohus, D., Horvitz, E., 2009 · 2009
Earlier work this paper cites.
3d is here: Point cloud library (pcl), in: 2011 IEEE international conference on robotics and automation, IEEE. pp. 1–4
Rusu, R.B., Cousins, S., 2011 · 2011
Earlier work this paper cites.
Long short-term memory
Graves, A., Graves, A., 2012 · 2012
Earlier work this paper cites.
Situated language understanding at 25 miles per hour, in: Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), pp. 22–31
Misu, T., Raux, A., Gupta, R., Lane, I., 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation, in: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543
Pennington, J., Socher, R., Manning, C.D., 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering, in: Proceedings of the IEEE international conference on computer vision, pp. 2425–2433
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D., 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575
Vedantam, R., Lawrence Zitnick, C., Parikh, D., 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation, in: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14, Springer. pp. 382–398
Anderson, P., Fernando, B., Johnson, M., Gould, S., 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text, in: Su, J., Duh, K., Carreras, X. (Eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Austin, Texas. pp. 2383–2392
Rajpurkar, P., Zhang, J., Lopyrev, K., Liang, P., 2016 · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5828–5839
Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Nießner, M., 2017 · 2017
Earlier work this paper cites.
A review of algorithms for filtering the 3d point cloud
Han, X.F., Jin, J.S., Wang, M.J., Jiang, W., Gao, L., Xiao, L., 2017 · 2017
Earlier work this paper cites.
Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 2961–2969
He, K., Gkioxari, G., Dollár, P., Girshick, R., 2017 · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2901–2910
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R., 2017 · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al., 2017 · 2017
Earlier work this paper cites.
Semantic scene completion from a single depth image, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1746–1754
Song, S., Yu, F., Zeng, A., Chang, A.X., Savva, M., Funkhouser, T., 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., 2017 · 2017
Earlier work this paper cites.
Embodied Question Answering, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D., 2018 · 2018
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4089–4098
Gordon, D., Kembhavi, A., Rastegari, M., Redmon, J., Fox, D., Farhadi, A., 2018 · 2018
Earlier work this paper cites.
Kudo, T., 2018 · 2018
Earlier work this paper cites.
Building generalizable agents with a realistic and rich 3d environment
Wu, Y., Wu, Y., Gkioxari, G., Tian, Y., 2018 · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 9068–9079
Xia, F., Zamir, A.R., He, Z., Sax, A., Malik, J., Savarese, S., 2018 · 2018
Earlier work this paper cites.
Resolving 3d human pose ambiguities with 3d scene constraints, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 2282–2292
Hassan, M., Choutas, V., Tzionas, D., Black, M.J., 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of naacL-HLT, Minneapolis, Minnesota. p. 2
Kenton, J.D.M.W.C., Toutanova, L.K., 2019 · 2019
Earlier work this paper cites.
Point-voxel cnn for efficient 3d deep learning
Liu, Z., Tang, H., Lin, Y., Han, S., 2019 · 2019
Earlier work this paper cites.
Deep hough voting for 3d object detection in point clouds, in: proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9277–9286
Qi, C.R., Litany, O., He, K., Guibas, L.J., 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al., 2019 · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 9339–9347
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al., 2019 · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7658–7667
Wald, J., Avetisyan, A., Navab, N., Tombari, F., Nießner, M., 2019 · 2019
Earlier work this paper cites.
Embodied question answering in photorealistic environments with point cloud perception, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6659–6668
Wijmans, E., Datta, S., Maksymets, O., Das, A., Gkioxari, G., Lee, S., Essa, I., Parikh, D., Batra, D., 2019 · 2019
Earlier work this paper cites.
Multi-target embodied question answering, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6309–6318
Yu, L., Chen, X., Gkioxari, G., Bansal, M., Berg, T.L., Batra, D., 2019 · 2019
Cited alongside, same era.
Information fusion in visual question answering: A survey
Zhang, D., Cao, R., Wu, S., 2019 · 2019
Cited alongside, same era.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer. pp. 422–440
Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., Guibas, L., 2020 · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language, in: European conference on computer vision, Springer. pp. 202–221
Chen, D.Z., Chang, A.X., Nießner, M., 2020 · 2020
Cited alongside, same era.
An embodied generalist agent in 3d world
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.C., Jia, B., Huang, S., 2023 · 2023
Later among the works it cites.
Frozen transformers in language models are effective visual encoder layers
Pang, Z., Xie, Z., Man, Y., Wang, Y.X., 2023 · 2023
Later among the works it cites.
Clip-guided vision-language pre-training for question answering in 3d scenes, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5607–5612
Parelli, M., Delitzas, A., Hars, N., Vlassis, G., Anagnostidis, S., Bachmann, G., Hofmann, T., 2023 · 2023
Later among the works it cites.
Openscene: 3d scene understanding with open vocabularies, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 815–824
Peng, S., Genova, K., Jiang, C., Tagliasacchi, A., Pollefeys, M., Funkhouser, T., et al., 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fang, H., Wang, C., Gou, M., Lu, C.G., · 2020
Cited alongside, same era.
Deep learning for 3d point clouds: A survey
Guo, Y., Wang, H., Hu, Q., Liu, H., Liu, L., Bennamoun, M., 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J., 2020 · 2020
Cited alongside, same era.
Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data
Baruch, G., Chen, Z., Dehghan, A., Dimry, T., Feigin, Y., Fu, P., Gebauer, T., Joffe, B., Kurz, D., Schwartz, A., et al., 2021 · 2021
Cited alongside, same era.
Scan2cap: Context-aware dense captioning in rgb-d scans, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3193–3203
Chen, Z., Gholami, A., Nießner, M., Chang, A.X., 2021 · 2021
Cited alongside, same era.
Stochastic scene-aware motion prediction, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11374–11384
Hassan, M., Ceylan, D., Villegas, R., Saito, J., Yang, J., Zhou, Y., Black, M.J., 2021 · 2021
Cited alongside, same era.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding, in: Proceedings of the 29th ACM International Conference on Multimedia, pp. 2344–2352
He, D., Zhao, Y., Luo, J., Hui, T., Huang, S., Zhang, A., Liu, S., 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR. pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Cited alongside, same era.
Mask3d: Mask transformer for 3d semantic instance segmentation, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), IEEE. pp. 8216–8223
Schult, J., Engelmann, F., Hermans, A., Litany, O., Tang, S., Leibe, B., 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al., 2023 · 2023
Later among the works it cites.
Habitat-matterport 3d semantics dataset, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4927–4936
Yadav, K., Ramrakhya, R., Ramakrishnan, S.K., Gervet, T., Turner, J., Gokaslan, A., Maestre, N., Chang, A.X., Batra, D., Savva, M., et al., 2023 · 2023
Later among the works it cites.
Comprehensive visual question answering on point clouds through compositional scene manipulation
Yan, X., Yuan, Z., Du, Y., Liao, Y., Guo, Y., Cui, S., Li, Z., 2023 · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.C., Liu, Z., Wang, L., 2023 · 2023
Later among the works it cites.
Uni3d: Exploring unified 3d representation at scale
Zhou, J., Wang, J., Ma, B., Liu, Y.S., Huang, T., Wang, X., 2023 · 2023
Later among the works it cites.
3d-vista: Pre-trained transformer for 3d vision and text alignment, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2911–2921
Zhu, Z., Ma, X., Chen, Y., Deng, Z., Huang, S., Li, Q., 2023 · 2023
Later among the works it cites.
Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26428–26438
Chen, S., Chen, X., Zhang, C., Li, M., Yu, G., Fei, H., Zhu, H., Fan, J., Chen, T., 2024 · 2024
Later among the works it cites.
Mmvqa: A comprehensive dataset for investigating multipage multimodal information retrieval in pdf-based visual question answering, in: 33rd International Joint Conference on Artificial Intelligence, IJCAI 2024, International Joint Conferences on Artificial Intelligence. pp. 6243–6251
Ding, Y., Ren, K., Huang, J., Luo, S., Han, S.C., 2024 · 2024
Later among the works it cites.
Scene-llm: Extending language model for 3d visual understanding and reasoning
Fu, R., Liu, J., Chen, X., Nie, Y., Xiong, W., 2024 · 2024
Later among the works it cites.
Chat-scene: Bridging 3d scene and large language models with object identifiers, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems
Huang, H., Chen, Y., Wang, Z., Huang, R., Xu, R., Wang, T., Liu, L., Cheng, X., Zhao, Y., Pang, J., et al., 2024 · 2024
Later among the works it cites.
From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities
Ishmam, M.F., Shovon, M.S.H., Mridha, M.F., Dey, N., 2024 · 2024
Later among the works it cites.
Multi-modal situated reasoning in 3d scenes
Linghu, X., Huang, J., Niu, X., Ma, X., Jia, B., Huang, S., 2024 · 2024
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16488–16498
Majumdar, A., Ajay, A., Zhang, X., Putta, P., Yenamandra, S., Henaff, M., Silwal, S., Mcvay, P., Maksymets, O., Arnaud, S., et al., 2024 · 2024
Later among the works it cites.
Situational awareness matters in 3d vision language reasoning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13678–13688
Man, Y., Gui, L.Y., Wang, Y.X., 2024 · 2024
Later among the works it cites.
Bridging the gap between 2d and 3d visual question answering: A fusion approach for 3d vqa, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 4261–4268
Mo, W., Liu, Y., 2024 · 2024
Later among the works it cites.
Langsplat: 3d language gaussian splatting, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20051–20060
Qin, M., Li, W., Zhou, J., Wang, H., Pfister, H., 2024 · 2024
Later among the works it cites.
Evaluating zero-shot gpt-4v performance on 3d visual question answering benchmarks
Singh, S., Pavlakos, G., Stamoulis, D., 2024 · 2024
Later among the works it cites.
Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics
Song, C.H., Blukis, V., Tremblay, J., Tyree, S., Su, Y., Birchfield, S., 2024 · 2024
Later among the works it cites.
Space3d-bench: Spatial 3d question answering benchmark
Szymanska, E., Dusmanu, M., Buurlage, J.W., Rad, M., Pollefeys, M., 2024 · 2024
Later among the works it cites.
Famma: A benchmark for financial domain multilingual multimodal question answering
Xue, S., Chen, T., Zhou, F., Dai, Q., Chu, Z., Mei, H., 2024 · 2024
Later among the works it cites.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Huang, X., Wang, Z., Sheng, L., Bai, L., et al., 2024 · 2024
Later among the works it cites.
Improved mlp point cloud processing with high-dimensional positional encoding, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 7891–7899
Zou, Y., Yu, H., Yang, Z., Li, Z., Akhtar, N., 2024 · 2024
Later among the works it cites.
Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery
Bai, L., Wang, G., Islam, M., Seenivasan, L., Wang, A., Ren, H., 2025 · 2025
Closest in time.
Deep learning based 3d segmentation in computer vision: A survey
He, Y., Yu, H., Liu, X., Yang, Z., Sun, W., Anwar, S., Mian, A., 2025 · 2025
Closest in time.
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding, in: European Conference on Computer Vision, Springer. pp. 289–310
Jia, B., Chen, Y., Yu, H., Wang, Y., Niu, X., Liu, T., Li, Q., Huang, S., 2025 · 2025
Closest in time.
From screens to scenes: A survey of embodied ai in healthcare
Liu, Y., Cao, X., Chen, T., Jiang, Y., You, J., Wu, M., Wang, X., Feng, M., Jin, Y., Chen, J., 2025 · 2025
Closest in time.
Splattalk: 3d vqa with gaussian splatting
Thai, A., Peng, S., Genova, K., Guibas, L., Funkhouser, T., 2025 · 2025
Closest in time.
Intervention and regulatory mechanism of multimodal fusion natural interactions on ar embodied cognition
Yong, J., Wei, J., Lei, X., Wang, Y., Dang, J., Lu, W., 2025 · 2025
Closest in time.
Empowering large language models with 3d situation awareness
Yuan, Z., Peng, Y., Ren, J., Liao, Y., Han, Y., Feng, C.M., Zhao, H., Li, G., Cui, S., Li, Z., 2025 · 2025
Closest in time.
His-gpt: Towards 3d human-in-scene multimodal understanding
Zhao, J., Hou, R., Tian, Z., Chang, H., Shan, S., 2025 · 2025
Closest in time.