Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) have become a pivotal research focus in deep learning, demonstrating remarkable capabilities in 3D scene understanding.
Goyal, Yash et al. “Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.” International Journal of Computer Vision 127 (2016): 398 - 414
2016
Earlier work this paper cites.
Dai, Angela et al. “ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes.” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017): 2432-2443
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
Chen, Dave Zhenyu et al. “Scan2Cap: Context-aware Dense Captioning in RGB-D Scans.” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020): 3192-3202
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Meng, Lingchen et al. “AdaViT: Adaptive Vision Transformers for Efficient Image Recognition.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 12299-12308
2021
Earlier work this paper cites.
Azuma, Daich et al. “ScanQA: 3D Question Answering for Spatial Scene Understanding.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 19107-19117
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Huang, Haifeng et al. “Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.” Neural Information Processing Systems (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Chen, Sijin et al. “LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning.” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023): 26418-26428
2023
Cited alongside, same era.
Liu, Haotian et al. “Visual Instruction Tuning.” ArXiv abs/2304.08485 (2023): n. pag
2023
Cited alongside, same era.
Bai, Jinze et al. “Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.” (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Liu, Yuanzhan et al. “MMBench: Is Your Multi-modal Model an All-around Player?” European Conference on Computer Vision (2023)
2023
Cited alongside, same era.
Achiam, OpenAI Josh et al. “GPT-4 Technical Report.” (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Zhang, Xiaofeng et al. “From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.” (2024)
2024
Cited alongside, same era.
Chen, Liang et al. “An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.” European Conference on Computer Vision (2024)
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.