Fetching the paper…
Reading the bibliography…
As large language models (LLMs) evolve, their integration with 3D spatial data (3D-LLMs) has seen rapid progress, offering unprecedented capabilities for understanding and interacting with physical spaces.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
A volumetric method for building complex models from range images
B. Curless and M. Levoy · 1996
Earlier work this paper cites.
A survey of augmented reality
R. T. Azuma · 1997
Earlier work this paper cites.
A touring machine: Prototyping 3d mobile augmented reality systems for exploring the urban environment
S. Feiner et al · 1997
Earlier work this paper cites.
A pde-based fast local level set method
D. Peng et al · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni et al · 2002
Earlier work this paper cites.
Level set methods and dynamic implicit surfaces
S. Osher et al · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Augmented reality: an overview
J. Carmigniani and B. Furht · 2011
Earlier work this paper cites.
Kinectfusion: Real-time dense surface mapping and tracking
R. A. Newcombe et al · 2011
Earlier work this paper cites.
Japanese and korean voice search
M. Schuster and K. Nakajima · 2012
Earlier work this paper cites.
Understanding augmented reality: Concepts and applications
A. B. Craig · 2013
Earlier work this paper cites.
Urban 3d semantic modelling using stereo vision
S. Sengupta et al · 2013
Earlier work this paper cites.
A Versatile Scene Model with Differentiable Visibility Applied to Generative Pose Estimation
H. Rhodin et al · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
R. Sennrich et al · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
R. Vedantam et al · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang et al · 2015
Earlier work this paper cites.
Structure-from-motion revisited
J. L. Schönberger and J.-M. Frahm · 2016
Earlier work this paper cites.
SPICE: Semantic propositional image caption evaluation
P. Anderson et al · 2016
Earlier work this paper cites.
3d semantic parsing of large-scale indoor spaces
I. Armeni et al · 2016
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi et al · 2017
Earlier work this paper cites.
Shape completion using 3d-encoder-predictor cnns and shape synthesis
A. Dai et al · 2017
Earlier work this paper cites.
Octnet: Learning deep 3d representations at high resolutions
G. Riegler et al · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
Semanticfusion: Dense 3d semantic mapping with convolutional neural networks
J. McCormac et al · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
A. Chang et al · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi et al · 2017
Earlier work this paper cites.
3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation
A. Dai and M. Nießner · 2018
Earlier work this paper cites.
Unsupervised Training for 3D Morphable Model Regression
K. Genova et al · 2018
Earlier work this paper cites.
Generating wikipedia by summarizing long sequences
P. J. Liu et al · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
J. Howard and S. Ruder · 2018
Earlier work this paper cites.
On evaluation of embodied navigation agents
P. Anderson et al · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson et al · 2018
Earlier work this paper cites.
Sgpn: Similarity group proposal network for 3d point cloud instance segmentation
W. Wang et al · 2018
Earlier work this paper cites.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Y. Zhou and O. Tuzel · 2018
Earlier work this paper cites.
Flownet3d: Learning scene flow in 3d point clouds
X. Liu et al · 2019
Earlier work this paper cites.
Learning view priors for single-view 3d reconstruction
H. Kato and T. Harada · 2019
Earlier work this paper cites.
Occupancy networks: Learning 3d reconstruction in function space
L. Mescheder et al · 2019
Earlier work this paper cites.
Deepsdf: Learning continuous signed distance functions for shape representation
J. J. Park et al · 2019
Earlier work this paper cites.
Scene representation networks: Continuous 3D-structure aware neural scene representations
V. Sitzmann et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
BERTScore: Evaluating text generation with bert
T. Zhang et al · 2019
Earlier work this paper cites.
3d-sis: 3d semantic instance segmentation of rgb-d scans
J. Hou et al · 2019
Earlier work this paper cites.
Apollocar3d: A large 3d car instance understanding benchmark for autonomous driving
X. Song et al · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments
J. Wald et al · 2019
Earlier work this paper cites.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
M. A. Uy et al · 2019
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
K. Chen et al · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
H. Caesar et al · 2019
Earlier work this paper cites.
Convolutional occupancy networks
S. Peng et al · 2020
Earlier work this paper cites.
NeRF Representing scenes as neural radiance fields for view synthesis
B. Mildenhall et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin et al · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen et al · 2020
Earlier work this paper cites.
Evaluation of text generation: A survey
A. Celikyilmaz et al · 2020
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
P. Achlioptas et al · 2020
Earlier work this paper cites.
Generative 3d part assembly via dynamic graph learning
J. Huang et al · 2020
Earlier work this paper cites.
Pointgroup: Dual-set point grouping for 3d instance segmentation
L. Jiang et al · 2020
Earlier work this paper cites.
Occuseg: Occupancy-aware 3d instance segmentation
L. Han et al · 2020
Earlier work this paper cites.
3d-aware scene change captioning from multiview images
Y. Qiu et al · 2020
Earlier work this paper cites.
Polygen: An autoregressive generative model of 3d meshes
C. Nash et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel et al · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford et al · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia et al · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron et al · 2021
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Z. Chen et al · 2021
Earlier work this paper cites.
Text-guided graph neural networks for referring 3d instance segmentation
P.-H. Huang et al · 2021
Earlier work this paper cites.
Holistic 3d scene understanding from a single image with implicit representation
C. Zhang et al · 2021
Earlier work this paper cites.
Deeppanocontext: Panoramic 3d scene understanding with holistic scene context graph and relation-based optimization
C. Zhang et al · 2021
Earlier work this paper cites.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
D. He et al · 2021
Earlier work this paper cites.
Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis
T. Shen et al · 2021
Earlier work this paper cites.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Z. Yuan et al · 2021
Earlier work this paper cites.
3dvg-transformer: Relation modeling for visual grounding on point clouds
L. Zhao et al · 2021
Earlier work this paper cites.
3d question answering, 2021
S. Ye et al · 2021
Earlier work this paper cites.
Both style and fog matter: Cumulative domain adaptation for semantic foggy scene understanding
X. Ma et al · 2022
Earlier work this paper cites.
Neural fields in visual computing and beyond
Y. Xie et al · 2022
Earlier work this paper cites.
Mip-nerf 360: Unbounded anti-aliased neural radiance fields
J. T. Barron et al · 2022
Earlier work this paper cites.
Plenoxels: Radiance fields without neural networks
A. Yu et al · 2022
Earlier work this paper cites.
Instant neural graphics primitives with a multiresolution hash encoding
T. Müller et al · 2022
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
D. Azuma et al · 2022
Earlier work this paper cites.
Sqa3d: Situated question answering in 3d scenes
X. Ma et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
J. Wei et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei et al · 2022
Earlier work this paper cites.
Regionclip: Region-based language-image pretraining
Y. Zhong et al · 2022
Earlier work this paper cites.
Image segmentation using text and image prompts
T. Lüddecke and A. Ecker · 2022
Earlier work this paper cites.
Expanding language-image pretrained models for general video recognition
B. Ni et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach et al · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho et al · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole et al · 2022
Earlier work this paper cites.
ibot: Image bert pre-training with online tokenizer
J. Zhou et al · 2022
Earlier work this paper cites.
Vision-and-language navigation: A survey of tasks, methods, and future directions
J. Gu et al · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar et al · 2022
Earlier work this paper cites.
A tri-layer plugin to improve occluded detection
G. Zhan et al · 2022
Earlier work this paper cites.
Leveraging large language models for robot 3d scene understanding
W. Chen et al · 2022
Earlier work this paper cites.
A systematic investigation of commonsense knowledge in large language models
X. L. Li et al · 2022
Earlier work this paper cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
X. Yu et al · 2022
Earlier work this paper cites.
Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them, 2022
R. Beaumont · 2022
Earlier work this paper cites.
Audioclip: Extending clip to image, text and audio
A. Guzhov et al · 2022
Earlier work this paper cites.
Slip: Self-supervision meets language-image pre-training
N. Mu et al · 2022
Cited alongside, same era.
Language-driven semantic segmentation
B. Li et al · 2022
Cited alongside, same era.
Scaling open-vocabulary image segmentation with image-level labels
G. Ghiasi et al · 2022
Cited alongside, same era.
Open-vocabulary queryable scene representations for real world planning
B. Chen et al · 2022
Cited alongside, same era.
Semantic Abstraction: Open-world 3D scene understanding from 2D vision-language models
H. Ha and S. Song · 2022
Cited alongside, same era.
Language-grounded indoor 3d semantic segmentation in the wild
D. Rozenberszki et al · 2022
Cited alongside, same era.
Multi-clip: Contrastive vision-language pre-training for question answering tasks in 3d scenes
A. Delitzas et al · 2023
Later among the works it cites.
Uni3dl: Unified model for 3d and language understanding
X. Li et al · 2023
Later among the works it cites.
Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
T. Wang et al · 2023
Later among the works it cites.
Z. Lin et al · 2023
Later among the works it cites.
Cross3dvg: Baseline and dataset for cross-dataset 3d visual grounding on different rgb-d scans
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decomposing nerf for editing via feature field distillation
S. Kobayashi et al · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach et al · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia et al · 2022
Cited alongside, same era.
Zero-shot text-guided object generation with dream fields
A. Jain et al · 2022
Cited alongside, same era.
Clip-mesh: Generating textured meshes from text using pretrained image-text models
N. M. Khalid et al · 2022
Cited alongside, same era.
Clip-forge: Towards zero-shot text-to-shape generation
A. Sanghi et al · 2022
Cited alongside, same era.
T. Miyanishi et al · 2023
Later among the works it cites.
Arkitscenerefer: Text-based localization of small objects in diverse real-world 3d indoor scenes
S. Kato et al · 2023
Later among the works it cites.
Comprehensive visual question answering on point clouds through compositional scene manipulation
X. Yan et al · 2023
Later among the works it cites.
M3dbench: Let’s instruct large models with multi-modal 3d prompts
M. Li et al · 2023
Later among the works it cites.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Z. Yin et al · 2023
Later among the works it cites.
Large language models for robotics: A survey
F. Zeng et al · 2023
Later among the works it cites.
Language-conditioned learning for robotic manipulation: A survey
H. Zhou et al · 2023
Later among the works it cites.
Aligning large language models with human: A survey, 2023
Y. Wang et al · 2023
Later among the works it cites.
Drive like a human: Rethinking autonomous driving with large language models
D. Fu et al · 2024
Closest in time.
Agent3d-zero: An agent for zero-shot 3d understanding, 2024
S. Zhang et al · 2024
Closest in time.
Multiply: A multisensory object-centric embodied large language model in 3d world
Y. Hong et al · 2024
Closest in time.
Gala3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting
X. Zhou et al · 2024
Closest in time.
Video understanding with large language models: A survey, 2024
Y. Tang et al · 2024
Closest in time.
Colmap-free 3d gaussian splatting
Y. Fu et al · 2024
Closest in time.
Instantsplat: Unbounded sparse-view pose-free gaussian splatting in 40 seconds
Z. Fan et al · 2024
Closest in time.
Large language models: A survey
S. Minaee et al · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
T. Dettmers et al · 2024
Closest in time.
Visual instruction tuning
H. Liu et al · 2024
Closest in time.
Emergent correspondence from image diffusion
L. Tang et al · 2024
Closest in time.
H.-H. Lee et al · 2024
Closest in time.
Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation
T. Wu et al · 2024
Closest in time.
Amodal ground truth and completion in the wild
G. Zhan et al · 2024
Closest in time.
Scenefun3d: Fine-grained functionality and affordance understanding in 3d scenes
A. Delitzas et al · 2024
Closest in time.
Scene-llm: Extending language model for 3d visual understanding and reasoning, 2024
R. Fu et al · 2024
Closest in time.
See, imagine, plan: Discovering and hallucinating tasks from a single image, 2024
C. Ma et al · 2024
Closest in time.
Situational awareness matters in 3d vision language reasoning
Y. Man et al · 2024
Closest in time.
An embodied generalist agent in 3d world
J. Huang et al · 2024
Closest in time.
Pointllm: Empowering large language models to understand point clouds
R. Xu et al · 2024
Closest in time.
Chat-scene: Bridging 3d scene and large language models with object identifiers
H. Huang et al · 2024
Closest in time.
3dmit: 3d multi-modal instruction tuning for scene understanding
Z. Li et al · 2024
Closest in time.
Shapellm: Universal 3d object understanding for embodied interaction
Z. Qi et al · 2024
Closest in time.
Minigpt-3d: Efficiently aligning 3d point clouds with large language models using 2d priors
Y. Tang et al · 2024
Closest in time.
Llana: Large language and nerf assistant
A. Amaduzzi et al · 2024
Closest in time.
More text, less point: Towards 3d data-efficient point-language understanding, 2024
Y. Tang et al · 2024
Closest in time.
Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness
C. Zhu et al · 2024
Closest in time.
3d-vla: A 3d vision-language-action generative world model
H. Zhen et al · 2024
Closest in time.
Holodeck: Language guided generation of 3d embodied ai environments
Y. Yang et al · 2024
Closest in time.
Llama-mesh: Unifying 3d mesh generation with language models
Z. Wang et al · 2024
Closest in time.
Spartun3d: Situated spatial understanding of 3d world in large language models
Y. Zhang et al · 2024
Closest in time.
Kestrel: Point grounding multimodal llm for part-aware 3d vision-language understanding, 2024
J. Fei et al · 2024
Closest in time.
SceneCraft: An LLM agent for synthesizing 3D scene as Blender code
Z. Hu et al · 2024
Closest in time.
Segment everything everywhere all at once
X. Zou et al · 2024
Closest in time.
Open-vocabulary 3d semantic segmentation with text-to-image diffusion models
X. Zhu et al · 2024
Closest in time.
Weakly supervised 3d open-vocabulary segmentation
K. Liu et al · 2024
Closest in time.
N2f2: Hierarchical scene understanding with nested neural feature fields
Y. Bhalgat et al · 2024
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Z. Wang et al · 2024
Closest in time.
Text2nerf: Text-driven 3d scene generation with neural radiance fields
J. Zhang et al · 2024
Closest in time.
Genzi: Zero-shot 3d human-scene interaction generation
L. Li and A. Dai · 2024
Closest in time.
Cg-hoi: Contact-guided 3d human-object interaction generation
C. Diller and A. Dai · 2024
Closest in time.
Graphdreamer: Compositional 3d scene synthesis from scene graphs
G. Gao et al · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
B. Chen et al · 2024
Closest in time.
Lexicon3d: Probing visual foundation models for complex 3d scene understanding
Y. Man et al · 2024
Closest in time.
Scalable 3d captioning with pretrained models
T. Luo et al · 2024
Closest in time.
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
B. Jia et al · 2024
Closest in time.
Scanents3d: Exploiting phrase-to-3d-object correspondences for improved visio-linguistic models in 3d scenes
A. Abdelreheem et al · 2024
Closest in time.
Multi-modal situated reasoning in 3d scenes
X. Linghu et al · 2024
Closest in time.
3d-grand: A million-scale dataset for 3d-llms with better grounding and less hallucination, 2024
J. Yang et al · 2024
Closest in time.
3d question answering for city scene understanding
P. Sun et al · 2024
Closest in time.
Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations
R. Lyu et al · 2024
Closest in time.
Grounded 3d-llm with referent tokens
Y. Chen et al · 2024
Closest in time.
Gaussian grouping: Segment and edit anything in 3d scenes
M. Ye et al · 2024
Closest in time.
Point transformer v3: Simpler, faster, stronger
X. Wu et al · 2024
Closest in time.
Affordancellm: Grounding affordance from vision language models
S. Qian et al · 2024
Closest in time.
Physically grounded vision-language models for robotic manipulation
J. Gao et al · 2024
Closest in time.
A general protocol to probe large vision models for 3d physical understanding
G. Zhan et al · 2024
Closest in time.
Behavior vision suite: Customizable dataset generation via simulation
Y. Ge et al · 2024
Closest in time.
Sceneteller: Language-to-3d scene generation
B. M. Öcal et al · 2024
Closest in time.
Scenecraft: Layout-guided 3d scene generation
X. Yang et al · 2024
Closest in time.
Hallucination of multimodal large language models: A survey, 2024
Z. Bai et al · 2024
Closest in time.
PhysDreamer: Physics-based interaction with 3d objects via video generation
T. Zhang et al · 2024
Closest in time.
Robin3d: Improving 3d large language model via robust instruction tuning, 2025
W. Kang et al · 2025
Closest in time.
Perla: Perceptive 3d language assistant
G. Mei et al · 2025
Closest in time.
Video-3d llm: Learning position-aware video representation for 3d scene understanding
D. Zheng et al · 2025
Closest in time.
Gpt4scene: Understand 3d scenes from videos with vision-language models
Z. Qi et al · 2025
Closest in time.
Lscenellm: Enhancing large 3d scene understanding using adaptive visual preferences
H. Zhi et al · 2025
Closest in time.
Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning
H. Yu et al · 2025
Closest in time.
Splattalk: 3d vqa with gaussian splatting
A. Thai et al · 2025
Closest in time.
Ross3d: Reconstructive visual instruction tuning with 3d-awareness
H. Wang et al · 2025
Closest in time.
3d-llava: Towards generalist 3d lmms with omni superpoint transformer
J. Deng et al · 2025
Closest in time.
Spatialllm: A compound 3d-informed design towards spatially-intelligent large multimodal models
W. Ma et al · 2025
Closest in time.
Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors, 2025
D. Zheng et al · 2025
Closest in time.
Vlm-3r: Vision-language models augmented with instruction-aligned 3d reconstruction
Z. Fan et al · 2025
Closest in time.
Leo-vl: Towards 3d vision-language generalists via data scaling with efficient representation
J. Huang et al · 2025
Closest in time.
Mllms need 3d-aware representation supervision for scene understanding
X. Huang et al · 2025
Closest in time.
3dllm-mem: Long-term spatial-temporal memory for embodied 3d large language model
W. Hu et al · 2025
Closest in time.
Revisiting 3d llm benchmarks: Are we really testing 3d capabilities?
J. Jin et al · 2025
Closest in time.
Vggt: Visual geometry grounded transformer
J. Wang et al · 2025
Closest in time.
Crossover: 3d scene cross-modal alignment
S. D. Sarkar et al · 2025
Closest in time.
Segment any 3d gaussians
J. Cen et al · 2025
Closest in time.
Integrating chain-of-thought for multimodal alignment: A study on 3d vision-language learning
Y. Chen et al · 2025
Closest in time.
Unveiling the mist over 3d vision-language understanding: Object-centric evaluation with chain-of-analysis
J. Huang et al · 2025
Closest in time.
From flatland to space: Teaching vision-language models to perceive and reason in 3d
J. Zhang et al · 2025
Closest in time.
Space3d-bench: Spatial 3d question answering benchmark
E. Szymańska et al · 2025
Closest in time.
Visual embodied brain: Let multimodal large language models see, think, and control in spaces
G. Luo et al · 2025
Closest in time.
Embodied intelligence for 3d understanding: A survey on 3d scene question answering
Z. Li et al · 2025
Closest in time.
Uniugg: Unified 3d understanding and generation via geometric-semantic encoding
Y. Xu et al · 2025
Closest in time.
Inferring dynamic physical properties from video foundation models, 2025
G. Zhan et al · 2025
Closest in time.
4d-bench: Benchmarking multi-modal large language models for 4d object understanding
W. Zhu et al · 2025
Closest in time.
Vlm4d: Towards spatiotemporal awareness in vision language models
S. Zhou et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo et al · 2025
Closest in time.