Fetching the paper…
Reading the bibliography…
We present LightVLA, a simple yet effective differentiable token pruning framework for vision-language-action (VLA) models.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Potapczynski, G. Loaiza-Ganem, and J. P. Cunningham, “Invertible gaussian reparameterization: Revisiting the gumbel-softmax,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 311–12 321, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer, “Learned token pruning for transformers,” in Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining , 2022, pp. 784–794
2022
Earlier work this paper cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. , “Lora: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th symposium on operating systems principles , 2023, pp. 611–626
2023
Earlier work this paper cites.
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” Advances in Neural Information Processing Systems , vol. 36, pp. 44 776–44 791, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF international conference on computer vision , 2023, pp. 11 975–11 986
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
C. Liang, S. Zhang, L. Zhang, Z. Wang, H. Wang, C.-L. Zhang, K.-W. Chang, T. Wang, and Y. You, “An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
T. Jiang, Q. Dong, Y. Ma, X. Ji, and Y. Liu, “Customizable multimodal trajectory prediction via nodes of interest selection for autonomous vehicles,” Expert Systems with Applications , vol. 288, p. 128222, 2025
2025
Closest in time.
2025
Closest in time.
Y. Yang, L. Zhang, Z. Chen, H. Wang, Y. Gao, Z. Wang, C.-L. Zhang, and K.-W. Chang, “Efficientvla: Training-free acceleration and compression for vision-language-action models,” 2025
2025
Closest in time.
S. Zhang, C.-L. Zhang, Z. Wang, and K.-W. Chang, “Llava-mini: Efficient image and video large multimodal models with one vision token,” 2025
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 26 296–26 306
2024
Cited alongside, same era.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 24 185–24 198
2024
Cited alongside, same era.
2024
Cited alongside, same era.
L. Zheng, L. Yin, Z. Xie, C. L. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, et al. , “Sglang: Efficient execution of structured language model programs,” Advances in neural information processing systems , vol. 37, pp. 62 557–62 583, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
Physical Intelligence, “ π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization,” 2025
2025
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Wen, Y. Zhu, J. Li, M. Zhu, Z. Tang, K. Wu, Z. Xu, N. Liu, R. Cheng, C. Shen, et al. , “Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation,” IEEE Robotics and Automation Letters , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
L. Cao, Z. Zhang, Y. Qu, and Y. Shen, “Fastvggt: Training-free acceleration of visual geometry transformer,” 2025
2025
Closest in time.