Fetching the paper…
Reading the bibliography…
In this report, we introduce Vintern-1B, a reliable 1-billion-parameters multimodal large language model (MLLM) for Vietnamese language tasks.
N. Nguyen, T. Nguyen, V. Tran, M.-T. Tran, T. D. Ngo, T. H. Nguyen, and M. Hoai, “Dictionary-guided scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7383–7392
2021
Earlier work this paper cites.
X.-S. Vu, Q.-A. Bui, N.-V. Nguyen, T. T. H. Nguyen, and T. Vu, “Mc-ocr challenge: Mobile-captured image document recognition for vietnamese receipts,” in 2021 RIVF International Conference on Computing and Communication Technologies (RIVF) . IEEE, 2021, pp. 1–6
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. H. Nguyen, D. T. Vo, and K. Van Nguyen, “Uit-hwdb: Using transferring method to construct a novel benchmark for evaluating unconstrained handwriting image recognition in vietnamese,” in 2022 RIVF International Conference on Computing and Communication Technologies (RIVF) . IEEE, 2022, pp. 659–664
2022
Earlier work this paper cites.
N. H. Nguyen, D. T. Vo, K. Van Nguyen, and N. L.-T. Nguyen, “Openvivqa: Task, dataset, and multimodal fusion models for visual question answering in vietnamese,” Information Fusion , vol. 100, p. 101868, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
“Gpt-4v(ision) system card,” 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:263218031
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , 2023
2023
Cited alongside, same era.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 24 185–24 198
2024
Closest in time.
2024
Closest in time.
B. Lin, Z. Tang, Y. Ye, J. Cui, B. Zhu, P. Jin, J. Zhang, M. Ning, and L. Yuan, “Moe-llava: Mixture of experts for large vision-language models,” 2024
2024
Closest in time.
H. V. Bui, H. H. Ha, P. V. Phan, and O. N. Tran, “Vistral v,” June 2024. [Online]. Available: https://huggingface.co/Vi-VLM/Vistral-V-7B
2024
Closest in time.
O. N. Tran, H. V. Bui, H. H. Ha, and P. V. Phan, “Vista,” May 2024. [Online]. Available: https://huggingface.co/datasets/Vi-VLM/Vista
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.