Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain blind to physical contact.
Measurement of shear and slip with a gelsight tactile sensor
Yuan, W., Li, R., Srinivasan, M. A., and Adelson, E. H · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gelsight: High-resolution robot tactile sensors for estimating geometry and force
Yuan, W., Dong, S., and Adelson, E. H · 2017
Earlier work this paper cites.
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Manipulation by feel: Touch-based control with deep predictive models
Tian, S., Ebert, F., Jayaraman, D., Mudigonda, M., Finn, C., Calandra, R., and Levine, S · 2019
Earlier work this paper cites.
Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation
Lambeta, M. et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Taxim: An example-based simulation model for gelsight tactile sensors
Si, Z. and Yuan, W · 2022
Earlier work this paper cites.
RT-1: Robotics Transformer for Real-World Control at Scale
Brohan, A. et al · 2023
Earlier work this paper cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Earlier work this paper cites.
See, hear, and feel: Smart sensory fusion for robotic manipulation
Li, H., Zhang, Y., Zhu, J., Wang, S., Lee, M. A., Xu, H., Adelson, E., Fei-Fei, L., Gao, R., and Wu, J · 2023
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2023
Earlier work this paper cites.
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B. et al · 2023
Earlier work this paper cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Black, K. et al · 2024
Cited alongside, same era.
A touch, vision, and language dataset for multimodal alignment
Fu, L. et al · 2024
Cited alongside, same era.
Octo: An Open-Source Generalist Robot Policy
Ghosh, D., Walke, H. R., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., Luo, J., Tan, Y. L., Chen, L. Y., Vuong, Q., Xiao, T., Sanketi, P. R., Sadigh, D., Finn, C., and Levine, S · 2024
Cited alongside, same era.
Li, Q., Liang, Y., Wang, Z., Luo, L., Chen, X., Liao, M., Wei, F., Deng, Y., Xu, S., Zhang, Y., et al · 2024
Cited alongside, same era.
Tacex: Gelsight tactile simulation in isaac sim – combining soft-body and visuotactile simulators
Nguyen, D. H., Duret, G., Schneider, T., Kshirsagar, A., Belousov, B., and Peters, J · 2024
Vitacformer: Learning cross-modal representation for visuo-tactile dexterous manipulation
Heng, L., Geng, H., Zhang, K., Abbeel, P., and Malik, J · 2025
Closest in time.
Sparsh: Self-supervised touch representations for vision-based tactile sensing
Higuera, C., Sharma, A., Bodduluri, C. K., Fan, T., Lancaster, P., Kalakrishnan, M., Kaess, M., Boots, B., Lambeta, M., Wu, T., and Mukadam, M · 2025
Closest in time.
Tactile-vla: Unlocking vision-language-action model’s physical knowledge for tactile generalization
Huang, J., Wang, S., Lin, F., Hu, Y., Wen, C., and Gao, Y · 2025
Closest in time.
Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding
Jones, J., Mees, O., Sferrazza, C., Stachowicz, K., Abbeel, P., and Levine, S · 2025
Closest in time.
Evo-0: Vision-language-action model with implicit spatial understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., et al · 2024
Cited alongside, same era.
Binding touch to everything: Learning unified multimodal tactile representations
Yang, F. et al · 2024
Cited alongside, same era.
V-jepa 2: Self-supervised video models enable understanding, prediction and planning
Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., Rizvi, A., Roberts, C., Sinha, K., Zholus, A., et al · 2025
Cited alongside, same era.
Anyskin: Plug-and-play skin sensing for robotic touch
Bhirangi, R., Pattabiraman, V., Erciyes, E., Cao, Y., Hellebrekers, T., and Pinto, L · 2025
Cited alongside, same era.
Gr00t n1: An open foundation model for generalist humanoid robots
Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L., Fang, Y., Fox, D., Hu, F., Huang, S., et al · 2025
Cited alongside, same era.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
Black, K. et al · 2025
Cited alongside, same era.
Worldvla: Towards autoregressive action world model
Cen, J., Yu, C., Yuan, H., Jiang, Y., Huang, S., Guo, J., Li, X., Song, Y., Luo, H., Wang, F., et al · 2025
Cited alongside, same era.
Lin, T., Li, G., Zhong, Y., Zou, Y., Du, Y., Liu, J., Gu, E., and Zhao, B · 2025
Closest in time.
Rdt-1b: a diffusion foundation model for bimanual manipulation
Liu, S., Wu, L., Li, B., Tan, H., Chen, H., Wang, Z., Xu, K., Su, H., and Zhu, J · 2025
Closest in time.
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models
Qu, D., Song, H., Chen, Q., Yao, Y., Ye, X., Gu, J., Wang, Z., Ding, Y., Zhao, B., Wang, D., and Li, X · 2025
Closest in time.
Vitacgen: Robotic pushing with vision-to-touch generation
Wu, Z., Lin, Y., Zhao, Y., Zhang, X., Chen, Z., Lepora, N., and Luo, S · 2025
Closest in time.
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
Xue, H., Ren, J., Chen, W., Zhang, G., Yuan, F., Gu, G., Xu, H., and Lu, C · 2025
Closest in time.
Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge
Zhang, W., Liu, H., Qi, Z., Wang, Y., Yu, X., Zhang, J., Dong, R., He, J., Wang, H., Zhang, Z., Yi, L., Zeng, W., and Jin, X · 2025
Closest in time.
Transferable tactile transformers for representation learning across diverse sensors and tasks
Zhao, J., Ma, Y., Wang, L., and Adelson, E · 2025
Closest in time.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Zheng, R., Liang, Y., Huang, S., Gao, J., Daumé III, H., Kolobov, A., Huang, F., and Yang, J · 2025
Closest in time.
Pointvla: Injecting the 3d world into vision-language-action models
Li, C., Wen, J., Peng, Y., Peng, Y., and Zhu, Y · 2026
Closest in time.