Fetching the paper…
Reading the bibliography…
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities - all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B).
“PyTorch Distributed: Experiences on Accelerating Data Parallel Training”, 2020
Shen Li et al · 2006
Earlier work this paper cites.
“Training Deep Nets with Sublinear Memory Cost”, 2016
Tianqi Chen et al · 2016
Earlier work this paper cites.
“A diagram is worth a dozen images”
Aniruddha Kembhavi et al · 2016
Earlier work this paper cites.
“GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism”, 2019
Yanping Huang et al · 2019
Earlier work this paper cites.
“Zero: Memory optimizations toward training trillion parameter models”
Samyam Rajbhandari et al · 2020
Earlier work this paper cites.
“Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM”, 2021
Deepak Narayanan et al · 2021
Earlier work this paper cites.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”, 2022
Tri Dao et al · 2022
Earlier work this paper cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”, 2022
William Fedus, Barret Zoph and Noam Shazeer · 2022
Earlier work this paper cites.
“Ego4d: Around the world in 3,000 hours of egocentric video”
Kristen Grauman et al · 2022
Earlier work this paper cites.
“Reducing Activation Recomputation in Large Transformer Models”, 2022
Vijay Korthikanti et al · 2022
Earlier work this paper cites.
“Infographicvqa”
Minesh Mathew et al · 2022
Earlier work this paper cites.
“Laion-5b: An open large-scale dataset for training next generation image-text models”
Christoph Schuhmann et al · 2022
Earlier work this paper cites.
“CoCa: Contrastive Captioners are Image-Text Foundation Models”, 2022
Jiahui Yu et al · 2022
Earlier work this paper cites.
“Amazon Simple Storage Service (Amazon S3)” Available at: https://aws.amazon.com/s3/ , Web, 2023
Amazon Web Services · 2023
Earlier work this paper cites.
“Patch n’ Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution”, 2023
Mostafa Dehghani et al · 2023
Earlier work this paper cites.
Sam Jacobs et al · 2023
Earlier work this paper cites.
“Ring Attention with Blockwise Transformers for Near-Infinite Context”, 2023
Hao Liu, Matei Zaharia and Pieter Abbeel · 2023
Earlier work this paper cites.
“MMBench: Is Your Multi-modal Model an All-around Player?”
Yuan Liu et al · 2023
Earlier work this paper cites.
“On the hidden mystery of ocr in large multimodal models”
Yuliang Liu et al · 2023
Earlier work this paper cites.
“Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts”
Pan Lu et al · 2023
Earlier work this paper cites.
“Egoschema: A diagnostic benchmark for very long-form video language understanding”
Karttikeya Mangalam, Raiymbek Akshulakov and Jitendra Malik · 2023
Earlier work this paper cites.
“RoFormer: Enhanced Transformer with Rotary Position Embedding”, 2023
Jianlin Su et al · 2023
Earlier work this paper cites.
“Mammoth: Building math generalist models through hybrid instruction tuning”
Xiang Yue et al · 2023
Cited alongside, same era.
“Sigmoid Loss for Language Image Pre-Training”, 2023
Xiaohua Zhai et al · 2023
Cited alongside, same era.
“Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale”, 2024
Rogerio Bonatti et al · 2024
Cited alongside, same era.
“Are We on the Right Way for Evaluating Large Vision-Language Models?”
Lin Chen et al · 2024
Cited alongside, same era.
“Seeclick: Harnessing gui grounding for advanced visual gui agents”
“Os-atlas: A foundation action model for generalist gui agents”
Zhiyong Wu et al · 2024
Later among the works it cites.
Zhiyu Wu et al · 2024
Later among the works it cites.
“Grok-1.5 Vision Preview”, 2024
x.ai · 2024
Later among the works it cites.
“Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments”
Tianbao Xie et al · 2024
Later among the works it cites.
“Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction”, 2024
Yiheng Xu et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kanzhi Cheng et al · 2024
Cited alongside, same era.
“Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis”
Chaoyou Fu et al · 2024
Cited alongside, same era.
“Blink: Multimodal large language models can see but not perceive”
Xingyu Fu et al · 2024
Cited alongside, same era.
“Datacomp: In search of the next generation of multimodal datasets”
Samir Gadre et al · 2024
Cited alongside, same era.
“MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale”, 2024
Jarvis Guo et al · 2024
Cited alongside, same era.
“Muon: An optimizer for hidden layers in neural networks”, 2024
Keller Jordan et al · 2024
Cited alongside, same era.
“Obelics: An open web-scale filtered dataset of interleaved image-text documents”
Hugo Laurençon et al · 2024
Cited alongside, same era.
“LLaVA-OneVision: Easy Visual Task Transfer”, 2024
Bo Li et al · 2024
Cited alongside, same era.
Jihan Yang et al · 2024
Later among the works it cites.
“Mm-vet: Evaluating large multimodal models for integrated capabilities”
Weihao Yu et al · 2024
Later among the works it cites.
“Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi”
Xiang Yue et al · 2024
Later among the works it cites.
“Mlvu: A comprehensive benchmark for multi-task long video understanding”
Junjie Zhou et al · 2024
Later among the works it cites.
“Multimodal c4: An open, billion-scale corpus of images interleaved with text”
Wanrong Zhu et al · 2024
Later among the works it cites.
“Qwen2.5-VL Technical Report”, 2025
Shuai Bai et al · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
DeepSeek-AI et al · 2025
Closest in time.
“DeepSeek-V3 Technical Report”, 2025
DeepSeek-AI et al · 2025
Closest in time.
“Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos”
Kairui Hu et al · 2025
Closest in time.
“ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use”
Kaixin Li et al · 2025
Closest in time.
“Muon is Scalable for LLM Training”
Jingyuan Liu et al · 2025
Closest in time.
“Muon is Scalable for LLM Training”
Jingyuan Liu et al · 2025
Closest in time.
“TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models”
Ziyao Shangguan et al · 2025
Closest in time.
“Gemma 3 Technical Report”, 2025
Gemma Team et al · 2025
Closest in time.
“Kimi k1. 5: Scaling reinforcement learning with llms”
Kimi Team et al · 2025
Closest in time.
“MMVU: Measuring Expert-Level Multi-Discipline Video Understanding”
Yilun Zhao et al · 2025
Closest in time.