Fetching the paper…
Reading the bibliography…
Recent advancements in reinforcement learning, particularly through Group Relative Policy Optimization (GRPO), have significantly improved multimodal large language models for complex reasoning tasks.
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec, “Trl: Transformer reinforcement learning,” https://github.com/huggingface/trl , 2020
2020
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
J. Xiao, A. Yao, Y. Li, and T. S. Chua, “Can i trust your answer? visually grounded video question answering,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
C. Yan, H. Wang, S. Yan, X. Jiang, Y. Hu, G. Kang, W. Xie, and E. Gavves, “Visa: Reasoning video object segmentation via large language models,” in European Conference on Computer Vision . Springer, 2024, pp. 98–115
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo et al. , “Mvbench: A comprehensive multi-modal video understanding benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 195–22 206
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Li, C. Wang, and J. Jia, “Llama-vid: An image is worth 2 tokens in large language models,” in European Conference on Computer Vision . Springer, 2024, pp. 323–340
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Cited alongside, same era.
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci, “R1-v: Reinforcing super generalization ability in vision-language models with less than $3,” https://github.com/Deep-Agent/R1-V , 2025, accessed: 2025-02-02
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.