Fetching the paper…
Reading the bibliography…
A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene.
Felzenszwalb, P.F., Huttenlocher, D.P.: Efficient graph-based image segmentation. International Journal of Computer Vision (IJCV) 59
2004
Earlier work this paper cites.
Qi, C.R., Su, H., Mo, K., Guibas, L.J.: Pointnet: Deep learning on point sets for 3d classification and segmentation. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 652–660 (2017)
2017
Earlier work this paper cites.
Qi, C.R., Yi, L., Su, H., Guibas, L.J.: Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems (NeurIPS) (2017)
2017
Earlier work this paper cites.
Choy, C., Gwak, J., Savarese, S.: 4d spatio-temporal convnets: Minkowski convolutional neural networks. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Earlier work this paper cites.
Ding, Z., Han, X., Niethammer, M.: Votenet: A deep learning label fusion method for multi-atlas segmentation. In: Medical Image Computing and Computer Assisted Intervention (MICCAI). Springer (2019)
2019
Earlier work this paper cites.
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al.: Habitat: A platform for embodied ai research. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 9339–9347 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., Guibas, L.: Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In: European Conference on Computer Vision (ECCV) (2020)
2020
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS) 33
2020
Earlier work this paper cites.
Chen, D.Z., Chang, A.X., Nießner, M.: Scanrefer: 3d object localization in rgb-d scans using natural language. In: European Conference on Computer Vision (ECCV) (2020)
2020
Earlier work this paper cites.
Jiang, L., Zhao, H., Shi, S., Liu, S., Fu, C.W., Jia, J.: Pointgroup: Dual-set point grouping for 3d instance segmentation. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Earlier work this paper cites.
Lin, Z., Zhang, Z., Chen, L.Z., Cheng, M.M., Lu, S.P.: Interactive image segmentation with first click attention. In: International Conference on Computer Vision (ICCV). pp. 13339–13348 (2020)
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21
2020
Earlier work this paper cites.
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable detr: Deformable transformers for end-to-end object detection. In: International Conference on Learning Representations (ICLR) (2020)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Chen, Z., Gholami, A., Nießner, M., Chang, A.X.: Scan2cap: Context-aware dense captioning in rgb-d scans. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Earlier work this paper cites.
Cho, J., Lei, J., Tan, H., Bansal, M.: Unifying vision-and-language tasks via text generation. In: International Conference on Machine Learning (ICML) (2021)
2021
Earlier work this paper cites.
Ding, H., Liu, C., Wang, S., Jiang, X.: Vision-language transformer and query generation for referring segmentation. In: International Conference on Computer Vision (ICCV). pp. 16321–16330 (2021)
2021
Earlier work this paper cites.
Dong, B., Zeng, F., Wang, T., Zhang, X., Wei, Y.: Solq: Segmenting objects by learning queries. Advances in Neural Information Processing Systems (NeurIPS) 34
2021
Earlier work this paper cites.
Fang, Y., Yang, S., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., Liu, W.: Instances as queries. In: International Conference on Computer Vision (ICCV). pp. 6910–6919 (2021)
2021
Earlier work this paper cites.
He, D., Zhao, Y., Luo, J., Hui, T., Huang, S., Zhang, A., Liu, S.: Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding. In: Proceedings of the 29th ACM International Conference on Multimedia (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Lei, J., Li, L., Zhou, L., Gan, Z., Berg, T.L., Bansal, M., Liu, J.: Less is more: Clipbert for video-and-language learning via sparse sampling. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021)
2021
Earlier work this paper cites.
Li, Y., Si, S., Li, G., Hsieh, C.J., Bengio, S.: Learnable fourier features for multi-dimensional spatial positional encoding. Advances in Neural Information Processing Systems (NeurIPS) 34
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (ICML). pp. 8748–8763. PMLR (2021)
2021
Earlier work this paper cites.
Yang, Z., Zhang, S., Wang, L., Luo, J.: Sat: 2d semantics assisted training for 3d visual grounding. In: International Conference on Computer Vision (ICCV) (2021)
2021
Earlier work this paper cites.
Zareian, A., Rosa, K.D., Hu, D.H., Chang, S.F.: Open-vocabulary object detection using captions. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 14393–14402 (2021)
2021
Cited alongside, same era.
Zhao, L., Cai, D., Sheng, L., Xu, D.: 3dvg-transformer: Relation modeling for visual grounding on point clouds. In: International Conference on Computer Vision (ICCV) (2021)
2021
Cited alongside, same era.
Azuma, D., Miyanishi, T., Kurita, S., Kawanabe, M.: Scanqa: 3d question answering for spatial scene understanding. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Cited alongside, same era.
Bakr, E.M., Alsaedy, Y., Elhoseiny, M.: Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding. Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Cited alongside, same era.
He, J., Li, P., Geng, Y., Xie, X.: Fastinst: A simple query-based model for real-time instance segmentation. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 23663–23672 (2023)
2023
Later among the works it cites.
Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., Gan, C.: 3d-llm: Injecting the 3d world into large language models. Advances in Neural Information Processing Systems (NeurIPS) 36
2023
Later among the works it cites.
Kerr, J., Kim, C.M., Goldberg, K., Kanazawa, A., Tancik, M.: Lerf: Language embedded radiance fields. In: International Conference on Computer Vision (ICCV). pp. 19729–19739 (2023)
2023
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: International Conference on Computer Vision (ICCV). pp. 4015–4026 (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., et al.: Do as i can, not as i say: Grounding language in robotic affordances. In: Conference on Robot Learning (CoRL) (2022)
2022
Cited alongside, same era.
Cai, D., Zhao, L., Zhang, J., Sheng, L., Xu, D.: 3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Cited alongside, same era.
Chen, S., Guhur, P.L., Tapaswi, M., Schmid, C., Laptev, I.: Language conditioned spatial relation reasoning for 3d object grounding. Advances in Neural Information Processing Systems (NeurIPS) (2022)
2022
Cited alongside, same era.
Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., Girdhar, R.: Masked-attention mask transformer for universal image segmentation. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Cited alongside, same era.
Duan, J., Yu, S., Tan, H.L., Zhu, H., Tan, C.: A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6
2022
Cited alongside, same era.
Ghiasi, G., Gu, X., Cui, Y., Lin, T.Y.: Scaling open-vocabulary image segmentation with image-level labels. In: European Conference on Computer Vision (ECCV). pp. 540–557. Springer (2022)
2022
Cited alongside, same era.
Ha, H., Song, S.: Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models. In: Conference on Robot Learning (CoRL) (2022)
2022
Cited alongside, same era.
Huang, S., Chen, Y., Jia, J., Wang, L.: Multi-view transformer for 3d visual grounding. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
Liang, F., Wu, B., Dai, X., Li, K., Zhao, Y., Zhang, H., Zhang, P., Vajda, P., Marculescu, D.: Open-vocabulary semantic segmentation with mask-adapted clip. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 7061–7070 (2023)
2023
Later among the works it cites.
Liu, Q., Xu, Z., Bertasius, G., Niethammer, M.: Simpleclick: Interactive image segmentation with simple vision transformers. In: International Conference on Computer Vision (ICCV). pp. 22290–22300 (2023)
2023
Later among the works it cites.
Ma, X., Yong, S., Zheng, Z., Li, Q., Liang, Y., Zhu, S.C., Huang, S.: Sqa3d: Situated question answering in 3d scenes. International Conference on Learning Representations (ICLR) (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Peng, S., Genova, K., Jiang, C., Tagliasacchi, A., Pollefeys, M., Funkhouser, T., et al.: Openscene: 3d scene understanding with open vocabularies. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 815–824 (2023)
2023
Later among the works it cites.
Qin, J., Wu, J., Yan, P., Li, M., Yuxi, R., Xiao, X., Wang, Y., Wang, R., Wen, S., Pan, X., et al.: Freeseg: Unified, universal and open-vocabulary image segmentation. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 19446–19455 (2023)
2023
Later among the works it cites.
Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I., Suenderhauf, N.: Sayplan: Grounding large language models using 3d scene graphs for scalable task planning. In: 7th Annual Conference on Robot Learning (2023)
2023
Later among the works it cites.
Schult, J., Engelmann, F., Hermans, A., Litany, O., Tang, S., Leibe, B.: Mask3d: Mask transformer for 3d semantic instance segmentation. In: International Conference on Robotics and Automation (ICRA). pp. 8216–8223. IEEE (2023)
2023
Later among the works it cites.
Takmaz, A., Fedele, E., Sumner, R.W., Pollefeys, M., Tombari, F., Engelmann, F.: OpenMask3D: Open-Vocabulary 3D Instance Segmentation. In: Advances in Neural Information Processing Systems (NeurIPS) (2023)
2023
Later among the works it cites.
Wu, Y., Cheng, X., Zhang, R., Cheng, Z., Zhang, J.: Eda: Explicit text-decoupling and dense alignment for 3d visual grounding. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 19231–19242 (2023)
2023
Later among the works it cites.
Xu, J., Hou, J., Zhang, Y., Feng, R., Wang, Y., Qiao, Y., Xie, W.: Learning open-vocabulary semantic segmentation models from natural language supervision. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 2935–2944 (2023)
2023
Later among the works it cites.
Xu, X., Xiong, T., Ding, Z., Tu, Z.: Masqclip for open-vocabulary universal image segmentation. In: International Conference on Computer Vision (ICCV). pp. 887–898 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhang, Y., Gong, Z., Chang, A.X.: Multi3drefer: Grounding text description to multiple 3d objects. In: International Conference on Computer Vision (ICCV). pp. 15225–15236 (2023)
2023
Later among the works it cites.
Zhu, Z., Ma, X., Chen, Y., Deng, Z., Huang, S., Li, Q.: 3d-vista: Pre-trained transformer for 3d vision and text alignment. In: International Conference on Computer Vision (ICCV). pp. 2911–2921 (2023)
2023
Later among the works it cites.
Zou, X., Dou, Z.Y., Yang, J., Gan, Z., Li, L., Li, C., Dai, X., Behl, H., Wang, J., Yuan, L., et al.: Generalized decoding for pixel, image, and language. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 15116–15127 (2023)
2023
Later among the works it cites.
Abdelreheem, A., Olszewski, K., Lee, H.Y., Wonka, P., Achlioptas, P.: Scanents3d: Exploiting phrase-to-3d-object correspondences for improved visio-linguistic models in 3d scenes. In: Proceedings of Winter Conference on Applications of Computer Vision (WACV). pp. 3524–3534 (2024)
2024
Closest in time.
Chen, S., Chen, X., Zhang, C., Li, M., Yu, G., Fei, H., Zhu, H., Fan, J., Chen, T.: Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning. In: The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024)
2024
Closest in time.
2024
Closest in time.
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.C., Jia, B., Huang, S.: An embodied generalist agent in 3d world. In: International Conference on Machine Learning (ICML) (2024)
2024
Closest in time.
Jia, B., Chen, Y., Yu, H., Wang, Y., Niu, X., Liu, T., Li, Q., Huang, S.: Sceneverse: Scaling 3d vision-language learning for grounded scene understanding. European Conference on Computer Vision (ECCV) (2024)
2024
Closest in time.