Fetching the paper…
Reading the bibliography…
3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments.
Synonyms provide semantic preview benefit in English
Elizabeth R Schotter. 2013 · 2013
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Tackling global grand challenges in our cities
Andrew Ka-Ching Chan. 2016 · 2016
Earlier work this paper cites.
Building functional cities
J Vernon Henderson, Anthony J Venables, Tanner Regan, and Ilia Samsonov. 2016 · 2016
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations
Thomas N Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5828–5839
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017 · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge. In Proceedings of the AAAI conference on artificial intelligence , Vol. 31
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Earlier work this paper cites.
Embodied question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1–10
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018 · 2018
Earlier work this paper cites.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
3-D scene graph: A sparse and semantic representation of physical environments for intelligent agents
Ue-Hwan Kim, Jin-Man Park, Taek-Jin Song, and Jong-Hwan Kim. 2019 · 2019
Earlier work this paper cites.
Deep hough voting for 3d object detection in point clouds. In proceedings of the IEEE/CVF International Conference on Computer Vision . 9277–9286
Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 3982–3992
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Embodied question answering in photorealistic environments with point cloud perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6659–6668
Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das, Georgia Gkioxari, Stefan Lee, Irfan Essa, Devi Parikh, and Dhruv Batra. 2019 · 2019
Earlier work this paper cites.
Multi-target embodied question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6309–6318
Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L Berg, and Dhruv Batra. 2019 · 2019
Earlier work this paper cites.
Real-time UAV path planning for autonomous urban scene reconstruction. In 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 1156–1162
Qi Kuang, Jinbo Wu, Jia Pan, and Bin Zhou. 2020 · 2020
Cited alongside, same era.
Moss: End-to-end dialog system framework with modular supervision. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 8327–8335
Weixin Liang, Youzhi Tian, Chengcai Chen, and Zhou Yu. 2020 · 2020
Cited alongside, same era.
A comparison of visual attention guiding approaches for 360 image-based vr tours. In 2020 IEEE Conference on Virtual Reality and 3D User Interfaces (VR) . IEEE, 83–91
Jan Oliver Wallgrün, Mahda M Bagher, Pejman Sajjadi, and Alexander Klippel. 2020 · 2020
Cited alongside, same era.
A survey of scene graph: Generation and application
Pengfei Xu, Xiaojun Chang, Ling Guo, Po-Yao Huang, Xiaojiang Chen, and Alexander G Hauptmann. 2020 · 2020
Cited alongside, same era.
Scan2cap: Context-aware dense captioning in rgb-d scans. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3193–3203
3D question answering
Shuquan Ye, Dongdong Chen, Songfang Han, and Jing Liao. 2022 · 2022
Later among the works it cites.
Towards Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline
Lichen Zhao, Daigang Cai, Jing Zhang, Lu Sheng, Dong Xu, Rui Zheng, Yinjie Zhao, Lipeng Wang, and Xibo Fan. 2022 · 2022
Later among the works it cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023a · 2023
Later among the works it cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023b · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhenyu Chen, Ali Gholami, Matthias Nießner, and Angel X Chang. 2021 · 2021
Cited alongside, same era.
A global horizon scan of the future impacts of robotics and autonomous systems on urban ecosystems
Mark A. Goddard, Zoe G. Davies, Solène Guenat, Mark J. Ferguson, Jessica C. Fisher, Adeniran Akanni, Teija Ahjokoski, Pippin M. L. Anderson, Fabio Angeoletto, Constantinos Antoniou, Adam J. Bates, Andrew Barkwith, Adam Berland, Christopher J. Bouch, Christine C. Rega-Brodsky, Loren B. Byrne, David Cameron, Rory Canavan, Tim Chapman, Stuart Connop, Steve Crossland, Marie C. Dade, David A. Dawson, Cynnamon Dobbs, Colleen T. Downs, Erle C. Ellis, Francisco J. Escobedo, Paul Gobster, Natalie Marie Gulsrud, Burak Guneralp, Amy K. Hahs, James D. Hale, Christopher Hassall, Marcus Hedblom, Dieter F. Hochuli, Tommi Inkinen, Ioan-Cristian Ioja, Dave Kendal, Tom Knowland, Ingo Kowarik, Simon J. Langdale, Susannah B. Lerman, Ian MacGregor-Fors, Peter Manning, Peter Massini, Stacey McLean, David D. Mkwambisi, Alessandro Ossola, Gabriel Pérez Luque, Luis Pérez-Urrestarazu, Katia Perini, Gad Perry, Tristan J. Pett, Kate E. Plummer, Raoufou A. Radji, Uri Roll, Simon G. Potts, Heather Rumble, Jon P. Sadler, Stevienna de Saille, Sebastian Sautter, Catherine E. Scott, Assaf Shwartz, Tracy Smith, Robbert P. H. Snep, Carl D. Soulsbury, Margaret C. Stanley, Tim Van de Voorde, Stephen J. Venn, Philip H. Warren, Carla-Leanne Washbourne, Mark Whitling, Nicholas S. G. Williams, Jun Yang, Kumelachew Yeshitela, Ken P. Yocom, and Martin Dallimer. 2021 · 2021
Cited alongside, same era.
Towards augmented reality driven human-city interaction: Current research on mobile headsets and future challenges
Lik-Hang Lee, Tristan Braud, Simo Hosio, and Pan Hui. 2021 · 2021
Cited alongside, same era.
Continuous aerial path planning for 3D urban scene reconstruction
Han Zhang, Yucong Yao, Ke Xie, Chi-Wing Fu, Hao Zhang, and Hui Huang. 2021 · 2021
Cited alongside, same era.
ScanQA: 3D question answering for spatial scene understanding. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 19129–19139
Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. 2022 · 2022
Cited alongside, same era.
Episodic memory question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19119–19128
Samyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai, Mukul Khanna, Dhruv Batra, and Devi Parikh. 2022 · 2022
Cited alongside, same era.
3dvqa: Visual question answering for 3d environments. In 2022 19th Conference on Robots and Vision (CRV) . IEEE, 233–240
Yasaman Etesam, Leon Kochiev, and Angel X Chang. 2022 · 2022
Cited alongside, same era.
Cric: A vqa dataset for compositional reasoning on vision and commonsense
Difei Gao, Ruiping Wang, Shiguang Shan, and Xilin Chen. 2022 · 2022
Cited alongside, same era.
3DGraphSeg: A unified graph representation-based point cloud segmentation framework for full-range highspeed railway environments
Yixuan Geng, Zhipeng Wang, Limin Jia, Yong Qin, Yuanyuan Chai, Keyan Liu, and Lei Tong. 2023 · 2023
Later among the works it cites.
Context-aware alignment and mutual masking for 3d-language pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10984–10994
Zhao Jin, Munawar Hayat, Yuwei Yang, Yulan Guo, and Yinjie Lei. 2023 · 2023
Later among the works it cites.
CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Taiki Miyanishi, Fumiya Kitamori, Shuhei Kurita, Jungdae Lee, Motoaki Kawanabe, and Nakamasa Inoue. 2023 · 2023
Later among the works it cites.
Clip-guided vision-language pre-training for question answering in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5606–5611
Maria Parelli, Alexandros Delitzas, Nikolas Hars, Georgios Vlassis, Sotirios Anagnostidis, Gregor Bachmann, and Thomas Hofmann. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
LLM-powered Data Augmentation for Enhanced Cross-lingual Performance. In The 2023 Conference on Empirical Methods in Natural Language Processing
Chenxi Whitehouse, Monojit Choudhury, and Alham Fikri Aji. 2023 · 2023
Later among the works it cites.
Comprehensive Visual Question Answering on Point Clouds through Compositional Scene Manipulation
Xu Yan, Zhihao Yuan, Yuhao Du, Yinghong Liao, Yao Guo, Shuguang Cui, and Zhen Li. 2023 · 2023
Later among the works it cites.
UrbanBIS: A Large-Scale Benchmark for Fine-Grained Urban Building Instance Segmentation. In ACM SIGGRAPH 2023 Conference Proceedings . 1–11
Guoqing Yang, Fuyou Xue, Qi Zhang, Ke Xie, Chi-Wing Fu, and Hui Huang. 2023 · 2023
Later among the works it cites.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions
Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny. 2023a · 2023
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 · 2024
Closest in time.
Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 4542–4550
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.