Fetching the paper…
Reading the bibliography…
Robotic chemists promise to both liberate human experts from repetitive tasks and accelerate scientific discovery, yet remain in their infancy.
Genetic k-means algorithm
K. Krishna and M. N. Murty · 1999
Earlier work this paper cites.
Fast marching farthest point sampling
C. Moenning and N. A. Dodgson · 2003
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al · 2018
Earlier work this paper cites.
Organic synthesis in a modular robotic system driven by a chemical programming language
S. Steiner, J. Wolf, S. Glatzel, A. Andreou, J. M. Granda, G. Keenan, T. Hinkley, G. Aragon-Camarasa, P. J. Kitson, D. Angelone, et al · 2019
Earlier work this paper cites.
A robotic platform for flow synthesis of organic compounds informed by ai planning
C. W. Coley, D. A. Thomas III, J. A. Lummiss, J. N. Jaworski, C. P. Breen, V. Schultz, T. Hart, J. S. Fishman, L. Rogers, H. Gao, et al · 2019
Earlier work this paper cites.
A mobile robotic chemist
B. Burger, P. M. Maffettone, V. V. Gusev, C. M. Aitchison, Y. Bai, X. Wang, X. Li, B. M. Alston, B. Li, R. Clowes, et al · 2020
Earlier work this paper cites.
A universal system for digitization and automatic execution of the chemical synthesis literature
S. H. M. Mehr, M. Craven, A. I. Leonov, G. Keenan, and L. Cronin · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Toist: Task oriented instance segmentation transformer with noun-pronoun distillation
P. Li, B. Tian, Y. Shi, X. Chen, H. Zhao, G. Zhou, and Y.-Q. Zhang · 2022
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Visual prompting via image inpainting
A. Bar, Y. Gandelsman, T. Darrell, A. Globerson, and A. Efros · 2022
Earlier work this paper cites.
Visual prompt tuning
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim · 2022
Earlier work this paper cites.
Chemistry lab automation via constrained task and motion planning
N. Yoshikawa, A. Z. Li, K. Darvish, Y. Zhao, H. Xu, A. Kuramshin, A. Aspuru-Guzik, A. Garg, and F. Shkurti · 2022
Earlier work this paper cites.
Archemist: Autonomous robotic chemistry system architecture
H. Fakhruldeen, G. Pizzuto, J. Glowacki, and A. I. Cooper · 2022
Earlier work this paper cites.
Core processes in intelligent robotic lab assistants: Flexible liquid handling
D. Knobbe, H. Zwirnmann, M. Eckhoff, and S. Haddadin · 2022
Earlier work this paper cites.
F-vlm: Open-vocabulary object detection upon frozen vision and language models
W. Kuo, Y. Cui, X. Gu, A. Piergiovanni, and A. Angelova · 2022
Earlier work this paper cites.
An autonomous laboratory for the accelerated synthesis of novel materials
N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gallant, E. D. Cubuk, A. Merchant, et al · 2023
Earlier work this paper cites.
Autonomous chemical research with large language models
D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes · 2023
Earlier work this paper cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Delving into shape-aware zero-shot semantic segmentation
X. Liu, B. Tian, Z. Wang, R. Wang, K. Sheng, B. Zhang, H. Zhao, and G. Zhou · 2023
Earlier work this paper cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Earlier work this paper cites.
Mvtrans: Multi-view perception of transparent objects
Y. R. Wang, Y. Zhao, H. Xu, S. Eppel, A. Aspuru-Guzik, F. Shkurti, and A. Garg · 2023
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T.-W. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Improving visual prompt tuning for self-supervised vision transformers
S. Yoo, E. Kim, D. Jung, J. Lee, and S. Yoon · 2023
Cited alongside, same era.
Explicit visual prompting for low-level structure segmentations
W. Liu, X. Shen, C.-M. Pun, and X. Cun · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Cited alongside, same era.
3D-VLA: A 3D vision-language-action generative world model
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan · 2024
Later among the works it cites.
Visual in-context prompting
F. Li, Q. Jiang, H. Zhang, T. Ren, S. Liu, X. Zou, H. Xu, H. Li, J. Yang, C. Li, et al · 2024
Later among the works it cites.
Vip-llava: Making large multimodal models understand arbitrary visual prompts
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y. Chai, D. Park, and Y. J. Lee · 2024
Later among the works it cites.
Moka: Open-world robotic manipulation through mark-based visual prompting
K. Fang, F. Liu, P. Abbeel, and S. Levine · 2024
Later among the works it cites.
Evaluation of microplate handling accuracy for applying robotic arms in laboratory automation
Y. Harazono, H. Shimono, K. Hata, T. Mitsuyama, and T. Horinouchi · 2024
Later among the works it cites.
Chemistry3d: Robotic interaction benchmark for chemistry experiments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Autonomous mobile robots for exploratory synthetic chemistry
T. Dai, S. Vijayakrishnan, F. T. Szczypiński, J.-F. Ayme, E. Simaei, T. Fellowes, R. Clowes, L. Kotopanov, C. E. Shields, Z. Zhou, et al · 2024
Cited alongside, same era.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Cited alongside, same era.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y. LeCun, and S. Xie · 2024
Cited alongside, same era.
Tod3cap: Towards 3d dense captioning in outdoor scenes
B. Jin, Y. Zheng, P. Li, W. Li, Y. Zheng, S. Hu, X. Liu, J. Zhu, Z. Yan, H. Sun, et al · 2024
Cited alongside, same era.
Hint-ad: Holistically aligned interpretability in end-to-end autonomous driving
K. Ding, B. Chen, Y. Su, H.-a. Gao, B. Jin, C. Sima, W. Zhang, X. Li, P. Barsch, H. Li, et al · 2024
Cited alongside, same era.
S. Li, Y. Huang, C. Guo, T. Wu, J. Zhang, L. Zhang, and W. Ding · 2024
Later among the works it cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
Impromptu vla: Open weights and open data for driving vision-language-action models
H. Chi, H.-a. Gao, Z. Liu, J. Liu, C. Liu, J. Li, K. Yang, Y. Yu, Z. Wang, W. Li, et al · 2025
Closest in time.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Chameleon: Fast-slow neuro-symbolic lane topology extraction
Z. Zhang, X. Li, S. Zou, G. Chi, S. Li, X. Qiu, G. Wang, G. Zheng, L. Wang, H. Zhao, et al · 2025
Closest in time.
AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, S. Jiang, Y. Jiang, C. Jing, H. Li, J. Li, C. Liu, Y. Liu, Y. Lu, J. Luo, P. Luo, Y. Mu, Y. Niu, Y. Pan, J. Pang, Y. Qiao, G. Ren, C. Ruan, J. Shan, Y. Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu, J. Yan, C. Yang, L. Yang, S. Yang, M. Yao, J. Zeng, C. Zhang, Q. Zhang, B. Zhao, C. Zhao, J. Zhao, and J. Zhu · 2025
Closest in time.
Universal actions for enhanced embodied foundation models
J. Zheng, J. Li, D. Liu, Y. Zheng, Z. Wang, Z. Ou, Y. Liu, J. Liu, Y.-Q. Zhang, and X. Zhan · 2025
Closest in time.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
J. Wen, Y. Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, et al · 2025
Closest in time.
Momanipvla: Transferring vision-language-action models for general mobile manipulation
Z. Wu, Y. Zhou, X. Xu, Z. Wang, and H. Yan · 2025
Closest in time.
Diffvla: Vision-language guided diffusion planning for autonomous driving
A. Jiang, Y. Gao, Z. Sun, Y. Wang, J. Wang, J. Chai, Q. Cao, Y. Heng, H. Jiang, Y. Dong, et al · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al · 2025
Closest in time.
Progressive visual prompt learning with contrastive feature re-formation
C. Xu, Y. Zhu, H. Shen, B. Chen, Y. Liao, X. Chen, and L. Wang · 2025
Closest in time.
Z. Liu, M. Zhang, and Y. Li · 2025
Closest in time.
Organa: a robotic assistant for automated chemistry experimentation and characterization
K. Darvish, M. Skreta, Y. Zhao, N. Yoshikawa, S. Som, M. Bogdanovic, Y. Cao, H. Hao, H. Xu, A. Aspuru-Guzik, et al · 2025
Closest in time.
Vision-based robot manipulation of transparent liquid containers in a laboratory setting
D. Schober, R. Güldenring, J. Love, and L. Nalpantidis · 2025
Closest in time.