Fetching the paper…
Reading the bibliography…
Building general-purpose intelligent robots has long been a fundamental goal of robotics.
Robot collisions: A survey on detection, isolation, and identification
S. Haddadin, A. De Luca, and A. Albu-Schäffer · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A review on cable-driven parallel robots
S. Qian, B. Zi, W.-W. Shang, and Q.-S. Xu · 2018
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Embodied intelligence via learning and evolution
A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Creating multimodal interactive agents with imitation and self-supervised learning
D. I. A. Team, J. Abramson, A. Ahuja, A. Brussee, F. Carnevale, M. Cassin, F. Fischer, P. Georgiev, A. Goldin, M. Gupta, et al · 2021
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. H. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. R. Florence · 2023
Cited alongside, same era.
Collision-free motion generation based on stochastic optimization and composite signed distance field networks of articulated robot
B. Liu, G. Jiang, F. Zhao, and X. Mei · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Real-time detailed self-collision avoidance in whole-body model predictive control
T. Jin, T. Kobayashi, and M. Doi · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Later among the works it cites.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al · 2023
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Learning with 3d rotations, a hitchhiker’s guide to so (3)
A. R. Geist, J. Frey, M. Zhobro, A. Levina, and G. Martius · 2024
Cited alongside, same era.
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Co-evolving embodied intelligence with design for artificial intelligence architecture
Y. Dai and J. Wang · 2025
Closest in time.
π _ 0.5 \pi\_0.5 : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Y. Jiang, R. Zhang, J. Wong, C. Wang, Y. Ze, H. Yin, C. Gokmen, S. Song, J. Wu, and L. Fei-Fei · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.