Fetching the paper…
Reading the bibliography…
The dawn of embodied intelligence has ushered in an unprecedented imperative for resilient, cognition-enabled multi-agent collaboration across next-generation ecosystems, revolutionizing paradigms in autonomous manufacturing, adaptive service robotics, and cyber-physical production architectures.
A tutorial on graph-based slam
G. Grisetti, R. Kümmerle, C. Stachniss, and W. Burgard · 2010
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Multi-robot cooperation
R. Fierro, L. Chaimowicz, and V. Kumar · 2018
Earlier work this paper cites.
Data-efficient multirobot, multitask transfer learning for trajectory tracking
K. Pereida, M. K. Helwa, and A. P. Schoellig · 2018
Earlier work this paper cites.
Cooperative heterogeneous multi-robot systems: A survey
Y. Rizk, M. Awad, and E. W. Tunstel · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments
J. Wald, A. Avetisyan, N. Navab, F. Tombari, and M. Nießner · 2019
Earlier work this paper cites.
Graph-structured referring expression reasoning in the wild
S. Yang, G. Li, and Y. Yu · 2020
Earlier work this paper cites.
Cops-ref: A new dataset and task on compositional referring expression comprehension
Z. Chen, P. Wang, L. Ma, K.-Y. K. Wong, and Q. Wu · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
A comprehensive survey of scene graphs: Generation and application
X. Chang, P. Ren, P. Xu, Z. Li, X. Chen, and A. Hauptmann · 2021
Earlier work this paper cites.
Scanqa: 3d question answering for spatial scene understanding
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe · 2022
Earlier work this paper cites.
Handmethat: Human-robot communication in physical and social environments
Y. Wan, J. Mao, and J. Tenenbaum · 2022
Earlier work this paper cites.
Multi-robot systems and cooperative object transport: Communications, platforms, and challenges
X. An, C. Wu, Y. Lin, M. Lin, T. Yoshinaga, and Y. Ji · 2023
Earlier work this paper cites.
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu · 2023
Earlier work this paper cites.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P. Luo · 2023
Earlier work this paper cites.
Vision-language foundation models as effective robot imitators
X. Li, M. Liu, H. Zhang, C. Yu, J. Xu, H. Wu, C. Cheang, Y. Jing, W. Zhang, H. Liu, et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Rtaw: An attention inspired reinforcement learning method for multi-robot task allocation in warehouse environments
A. Agrawal, A. S. Bedi, and D. Manocha · 2023
Earlier work this paper cites.
Cross-entropy regularized policy gradient for multirobot nonadversarial moving target search
H. Guo, Z. Liu, R. Shi, W.-Y. Yau, and D. Rus · 2023
Earlier work this paper cites.
On collaborative robot teams for environmental monitoring: A macroscopic ensemble approach
V. Edwards, T. C. Silva, B. Mehta, J. Dhanoa, and M. A. Hsieh · 2023
Earlier work this paper cites.
Learning to navigate in turbulent flows with aerial robot swarms: A cooperative deep reinforcement learning approach
D. Patiño, S. Mayya, J. Calderon, K. Daniilidis, and D. Saldaña · 2023
Earlier work this paper cites.
How to guide your learner: Imitation learning with active adaptive expert involvement
X.-H. Liu, F. Xu, X. Zhang, T. Liu, S. Jiang, R. Chen, Z. Zhang, and Y. Yu · 2023
Earlier work this paper cites.
Mitigating hallucination in large multi-modal models via robust instruction tuning
F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, and L. Wang · 2023
Earlier work this paper cites.
Sqa3d: Situated question answering in 3d scenes
X. Ma, S. Yong, Z. Zheng, Q. Li, Y. Liang, S.-C. Zhu, and S. Huang · 2023
Earlier work this paper cites.
Paco: Parts and attributes of common objects
V. Ramanathan, A. Kalia, V. Petrovic, Y. Wen, B. Zheng, B. Guo, R. Wang, A. Marquez, R. Kovvuri, A. Kadian, et al · 2023
Earlier work this paper cites.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Earlier work this paper cites.
Roco: Dialectic multi-robot collaboration with large language models
Z. Mandi, S. Jain, and S. Song · 2024
Earlier work this paper cites.
Coherent: Collaboration of heterogeneous multi-robot system with large language models
K. Liu, Z. Tang, D. Wang, Z. Wang, X. Li, and B. Zhao · 2024
Earlier work this paper cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Cited alongside, same era.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Cited alongside, same era.
π _ 0 \pi\_0 : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Navid: Video-based vlm plans the next step for vision-and-language navigation
J. Zhang, K. Wang, R. Xu, G. Zhou, Y. Hong, X. Fang, Q. Wu, Z. Zhang, and H. Wang · 2024
Cited alongside, same era.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Hi robot: Open-ended instruction following with hierarchical vision-language-action models
L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tanner, A. Walling, H. Wang, N. Fusai, et al · 2025
Closest in time.
π \pi 0.5: A vision-language-action model with open-world generalization
Physical Intelligence · 2025
Closest in time.
Cerebellar output shapes cortical preparatory activity during motor adaptation
S. Israely, H. Ninou, O. Rajchert, L. Elmaleh, R. Harel, F. Mawase, J. Kadmon, and Y. Prut · 2025
Closest in time.
Neuronal dynamics of cerebellum and medial prefrontal cortex in adaptive motor timing
Z. Ren, X. Wang, M. Angelov, C. I. De Zeeuw, and Z. Gao · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al · 2024
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl, et al · 2024
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al · 2024
Cited alongside, same era.
What foundation models can bring for robot learning in manipulation: A survey
D. Li, Y. Jin, Y. Sun, H. Yu, J. Shi, X. Hao, P. Hao, H. Liu, F. Sun, J. Zhang, et al · 2024
Cited alongside, same era.
Robomamba: Multimodal state space model for efficient robot reasoning and manipulation
J. Liu, M. Liu, Z. Wang, L. Lee, K. Zhou, P. An, S. Yang, R. Zhang, Y. Guo, and S. Zhang · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Cited alongside, same era.
Closest in time.
Dual and plasticity-dependent regulation of cerebello-zona incerta circuits on anxiety-like behaviors
Y. Zhao, J.-T. Wu, J.-B. Feng, X.-Y. Cai, X.-T. Wang, L. Wang, W. Xie, Y. Gu, J. Liu, W. Chen, et al · 2025
Closest in time.
Robobrain: A unified brain model for robotic manipulation from abstract to concrete
Y. Ji, H. Tan, J. Shi, X. Hao, Y. Zhang, H. Zhang, P. Wang, M. Zhao, Y. Mu, P. An, et al · 2025
Closest in time.
Affordgrasp: In-context affordance reasoning for open-vocabulary task-oriented grasping in clutter
Y. Tang, S. Zhang, X. Hao, P. Wang, J. Wu, Z. Wang, and S. Zhang · 2025
Closest in time.
L. Zhang, X. Hao, Q. Xu, Q. Zhang, X. Zhang, P. Wang, J. Zhang, Z. Wang, S. Zhang, and R. Xu · 2025
Closest in time.
FAST_LIO_LOCALIZATION_HUMANOID
Y. Xu, X. Li, Z. Zhao, G. Zhou, D. Li, X. An, H. Tan, and Z. Feng · 2025
Closest in time.
Cordvip: Correspondence-based visuomotor policy for dexterous manipulation in real-world
Y. Fu, Q. Feng, N. Chen, Z. Zhou, M. Liu, M. Wu, T. Chen, S. Rong, J. Liu, H. Dong, et al · 2025
Closest in time.
Dexgrasp anything: Towards universal robotic dexterous grasping with physics awareness
Y. Zhong, Q. Jiang, J. Yu, and Y. Ma · 2025
Closest in time.
Planning and control for deformable linear object manipulation
B. Aksoy and J. Wen · 2025
Closest in time.
Flagscale
BAAI · 2025
Closest in time.
Introducing claude 3.5 sonnet
Anthropic · 2025
Closest in time.
Evaluating gpt-4o’s embodied intelligence: A comprehensive empirical study
Y. Wu, H. Lyu, Y. Tang, L. Zhang, Z. Zhang, W. Zhou, and S. Hao · 2025
Closest in time.
Introducing gemini: Our largest and most capable ai model
Google · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
Learning to reason with llms
OpenAI · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Reason-rft: Reinforcement fine-tuning for visual reasoning
H. Tan, Y. Ji, X. Hao, M. Lin, P. Wang, Z. Wang, and S. Zhang · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Y. Hu, and S. Lin · 2025
Closest in time.
Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks
W. Zhang, M. Wang, G. Liu, X. Huixin, Y. Jiang, Y. Shen, G. Hou, Z. Zheng, H. Zhang, X. Li, et al · 2025
Closest in time.
Tla: Tactile-language-action model for contact-rich manipulation
P. Hao, C. Zhang, D. Li, X. Cao, X. Hao, S. Cui, and S. Wang · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al · 2025
Closest in time.
Mapfusion: A novel bev feature fusion network for multi-modal map construction
X. Hao, Y. Diao, M. Wei, Y. Yang, P. Hao, R. Yin, H. Zhang, W. Li, S. Zhao, and Y. Liu · 2025
Closest in time.
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Huang, S. Jiang, et al · 2025
Closest in time.
Qwen3: Think deeper, act faster
Q. Team · 2025
Closest in time.