Fetching the paper…
Reading the bibliography…
Latent Action Models (LAMs) have rapidly gained traction as an important component in the pre-training pipelines of leading Vision-Language-Action models.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Pont-Tuset, J., Perazzi, F., Caelles, S., Arbeláez, P., Sorkine-Hornung, A., and Van Gool, L · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Liu, Y., Gupta, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Learning what you can do before doing anything
Rybkin, O., Pertsch, K., Derpanis, K. G., Daniilidis, K., and Jaegle, A · 2018
Earlier work this paper cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Earlier work this paper cites.
Recent advances in imitation learning from observation
Torabi, F., Warnell, G., and Stone, P · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Earlier work this paper cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2023
Earlier work this paper cites.
Learning to act without actions
Schmidt, D. and Jiang, M · 2023
Cited alongside, same era.
Semi-supervised offline reinforcement learning with action-free trajectories
Zheng, Q., Henaff, M., Amos, B., and Grover, A · 2023
Cited alongside, same era.
Agrawal, P., Antoniak, S., Hanna, E. B., Bout, B., Chaplot, D., Chudnovsky, J., Costa, D., De Monicault, B., Garg, S., Gervet, T., et al · 2024
Cited alongside, same era.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Cited alongside, same era.
Dynamo: In-domain dynamics pretraining for visuo-motor control
Cui, Z. J., Pan, H., Iyer, A., Haldar, S., and Pinto, L · 2024
Adaworld: Learning adaptable world models with latent actions
Gao, S., Zhou, S., Du, Y., Zhang, J., and Gan, C · 2025
Later among the works it cites.
Breaking the modality barrier: Universal embedding learning with multimodal llms
Gu, T., Yang, K., Feng, Z., Wang, X., Zhang, Y., Long, D., Chen, Y., Cai, W., and Deng, J · 2025
Later among the works it cites.
Otter: A vision-language-action model with text-aware visual feature extraction
Huang, H., Liu, F., Fu, L., Wu, T., Mukadam, M., Malik, J., Goldberg, K., and Abbeel, P · 2025
Later among the works it cites.
Dreamgen: Unlocking generalization in robot learning through video world models
Jang, J., Ye, S., Lin, Z., Xiang, J., Bjorck, J., Fang, Y., Hu, F., Huang, S., Kundalia, K., Lin, Y.-C., et al · 2025
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How well can vision language models see image details?
Gou, C., Felemban, A., Khan, F. F., Zhu, D., Cai, J., Rezatofighi, H., and Elhoseiny, M · 2024
Cited alongside, same era.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., et al · 2024
Cited alongside, same era.
Vision language models are blind: Failing to translate detailed visual features into words
Rahmanzadehgervi, P., Bolton, L., Taesiri, M. R., and Nguyen, A. T · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al · 2024
Cited alongside, same era.
Latent action pretraining from videos
Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al · 2024
Cited alongside, same era.
Cosmos-reason1: From physical common sense to embodied reasoning
Azzolini, A., Bai, J., Brandon, H., Cao, J., Chattopadhyay, P., Chen, H., Chu, J., Cui, Y., Diamond, J., Ding, Y., et al · 2025
Cited alongside, same era.
Klepach, A., Nikulin, A., Zisman, I., Tarasov, D., Derevyagin, A., Polubarov, A., Lyubaykin, N., and Kurenkov, V · 2025
Later among the works it cites.
Lost in embeddings: Information loss in vision-language models
Li, W., Tang, R., Li, C., Zhang, C., Vulić, I., and Søgaard, A · 2025
Later among the works it cites.
Clam: Continuous latent action models for robot learning from unlabeled demonstrations
Liang, A., Czempin, P., Hong, M., Zhou, Y., Biyik, E., and Tu, S · 2025
Later among the works it cites.
Towards generalist robot learning from internet video: A survey
McCarthy, R., Tan, D. C., Schmidt, D., Acero, F., Herr, N., Du, Y., Thuruthel, T. G., and Li, Z · 2025
Later among the works it cites.
Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents
Meng, R., Jiang, Z., Liu, Y., Su, M., Yang, X., Fu, Y., Qin, C., Chen, Z., Xu, R., Xiong, C., et al · 2025
Later among the works it cites.
Latent action learning requires supervision in the presence of distractors
Nikulin, A., Zisman, I., Tarasov, D., Lyubaykin, N., Polubarov, A., Kiselev, I., and Kurenkov, V · 2025
Later among the works it cites.
Can VLMs actually see and read? a survey on modality collapse in vision-language models
Sim, M. Y., Zhang, W. E., Dai, X., and Fang, B · 2025
Later among the works it cites.
Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ramé, A., Rivière, M., et al · 2025
Later among the works it cites.
Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos
Yuan, H., Li, X., Zhang, T., Huang, Z., Xu, S., Ji, S., Tong, Y., Qi, L., Feng, J., and Yang, M.-H · 2025
Later among the works it cites.
What do latent action models actually learn?
Zhang, C., Pearce, T., Zhang, P., Wang, K., Chen, X., Shen, W., Zhao, L., and Bian, J · 2025
Later among the works it cites.
A survey on vision-language-action models: An action tokenization perspective
Zhong, Y., Bai, F., Cai, S., Huang, X., Chen, Z., Zhang, X., Wang, Y., Guo, S., Guan, T., Lui, K. N., et al · 2025
Later among the works it cites.