Fetching the paper…
Reading the bibliography…
We present Megrez2, a novel lightweight and high-performance language model architecture optimized for device native deployment.
“High-dimensional continuous control using generalized advantage estimation”
John Schulman et al · 2015
Earlier work this paper cites.
“Program synthesis with large language models”
Jacob Austin et al · 2021
Earlier work this paper cites.
“Evaluating Large Language Models Trained on Code”
Mark Chen et al · 2021
Earlier work this paper cites.
“Training Verifiers to Solve Math Word Problems”
Karl Cobbe et al · 2021
Earlier work this paper cites.
“Measuring mathematical problem solving with the math dataset”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
Josh Achiam et al · 2023
Earlier work this paper cites.
“Fast inference of mixture-of-experts language models with offloading”
Artyom Eliseev and Denis Mazur · 2023
Earlier work this paper cites.
“Instruction-following evaluation for large language models”
Jeffrey Zhou et al · 2023
Earlier work this paper cites.
“ Read-ME
Ruisi Cai et al · 2024
Earlier work this paper cites.
“Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models”
Damai Dai et al · 2024
Earlier work this paper cites.
“C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models”
Yuzhen Huang et al · 2024
Cited alongside, same era.
“Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference”
Ranggi Hwang et al · 2024
Cited alongside, same era.
“Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model”
Aixin Liu et al · 2024
Cited alongside, same era.
“GPT-4o mini: advancing cost-efficient intelligence” Released July 18, 2024, OpenAI blog post, 2024
OpenAI · 2024
Cited alongside, same era.
“Deepseekmath: Pushing the limits of mathematical reasoning in open language models”
Zhihong Shao et al · 2024
Cited alongside, same era.
“Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras”
Abdelrahman Abouelenin et al · 2025
Closest in time.
“Kimi-K2: Open-Source Models by Moonshot AI”, https://github.com/MoonshotAI/Kimi-K2 , 2025
Moonshot AI · 2025
Closest in time.
“Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning”
Daya Guo et al · 2025
Closest in time.
Gemma Kamath et al · 2025
Closest in time.
“Megrez-omni technical report”
Boxun Li et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Cited alongside, same era.
“Auxiliary-loss-free load balancing strategy for mixture-of-experts”
Lean Wang et al · 2024
Cited alongside, same era.
“Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024”
Yubo Wang et al · 2024
Cited alongside, same era.
“Skywork-moe: A deep dive into training techniques for mixture-of-experts language models”
Tianwen Wei et al · 2024
Cited alongside, same era.
Qwen Yang et al · 2024
Cited alongside, same era.
“The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation”, https://ai.meta.com/blog/llama-4-multimodal-intelligence/ , 2025
Meta AI · 2025
Closest in time.
“Introducing OpenAI o3 and o4-mini”, 2025
OpenAI · 2025
Closest in time.
“Gemini: A Family of Highly Capable Multimodal Models”, 2025
Gemini Team et al · 2025
Closest in time.
An Yang et al · 2025
Closest in time.