Fetching the paper…
Reading the bibliography…
With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI agents toward more dynamic, proactive, and generalizable agentic AI.
Handwritten digit recognition with a back-propagation network
Y. LeCun et al · 1989
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
R. Coulom · 2006
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman et al · 2017
Earlier work this paper cites.
Neural modular control for embodied question answering
A. Das et al · 2018
Earlier work this paper cites.
Learning to look around: Intelligently exploring unseen environments for unknown tasks
D. Jayaraman and K. Grauman · 2018
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi et al · 2019
Earlier work this paper cites.
Learning to explore using active neural slam
D. S. Chaplot et al · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra et al · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen et al · 2021
Earlier work this paper cites.
Safe reinforcement learning using formal verification for tissue retraction in autonomous robotic-assisted surgery
A. Pore et al · 2021
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
S. Yao et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima et al · 2022
Earlier work this paper cites.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
W. Chen et al · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Y. Li et al · 2022
Earlier work this paper cites.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
H. Le et al · 2022
Earlier work this paper cites.
Gpt4vis: What can gpt-4 do for zero-shot visual recognition?
W. Wu et al · 2023
Earlier work this paper cites.
Llm4drive: A survey of large language models for autonomous driving
Z. Yang et al · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
X. Deng et al · 2023
Earlier work this paper cites.
Memorybank: Enhancing large language models with long-term memory, 2023
W. Zhong et al · 2023
Earlier work this paper cites.
Tora: A tool-integrated reasoning agent for mathematical problem solving
Z. Gou et al · 2023
Earlier work this paper cites.
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning
K. Wang et al · 2023
Earlier work this paper cites.
Alp: Action-aware embodied learning for perception
X. Liang et al · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li et al · 2023
Earlier work this paper cites.
Valley: Video assistant with large language model enhanced ability
R. Luo et al · 2023
Earlier work this paper cites.
Visual instruction tuning
H. Liu et al · 2023
Earlier work this paper cites.
Video-llava: Learning united visual representation by alignment before projection
B. Lin et al · 2023
Earlier work this paper cites.
J. Ye et al · 2023
Earlier work this paper cites.
Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts
Y. Li et al · 2023
Earlier work this paper cites.
Continual pre-training of language models
Z. Ke et al · 2023
Earlier work this paper cites.
Continual pre-training of large language models: How to (re) warm your model?
K. Gupta et al · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao et al · 2023
Earlier work this paper cites.
Metatool benchmark for large language models: Deciding whether to use tools and which to use
Y. Huang et al · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu et al · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao et al · 2023
Earlier work this paper cites.
Multimodal chain-of-thought reasoning in language models
Z. Zhang et al · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn et al · 2023
Earlier work this paper cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
H. Zhang et al · 2023
Earlier work this paper cites.
Memgpt: Towards llms as operating systems
C. Packer et al · 2023
Earlier work this paper cites.
Avis: Autonomous visual information seeking with large language models
Z. Hu et al · 2023
Earlier work this paper cites.
Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory
Z. Hu et al · 2023
Earlier work this paper cites.
Execution-based code generation using deep reinforcement learning
P. Shojaee et al · 2023
Earlier work this paper cites.
Segment anything, 2023
A. Kirillov et al · 2023
Earlier work this paper cites.
Unsloth, 2023
M. H. Daniel Han and U. team · 2023
Earlier work this paper cites.
Fireact: Toward language agent fine-tuning, 2023
B. Chen et al · 2023
Earlier work this paper cites.
Agenttuning: Enabling generalized agent abilities for llms, 2023
A. Zeng et al · 2023
Earlier work this paper cites.
Lmflow: An extensible toolkit for finetuning and inference of large foundation models
S. Diao et al · 2023
Earlier work this paper cites.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
C. Li et al · 2023
Earlier work this paper cites.
Huatuogpt, towards taming language model to be a doctor
H. Zhang et al · 2023
Earlier work this paper cites.
Huatuogpt-ii, one-stage training for medical adaption of llms
J. Chen et al · 2023
Earlier work this paper cites.
Robotic-assisted navigation system for preoperative lung nodule localization: a pilot study
J. Liu et al · 2023
Earlier work this paper cites.
Minicpm-v: A gpt-4v level mllm on your phone
Y. Yao et al · 2024
Earlier work this paper cites.
Next-gpt: Any-to-any multimodal llm
S. Wu et al · 2024
Earlier work this paper cites.
Anygpt: Unified multimodal llm with discrete sequence modeling
J. Zhan et al · 2024
Earlier work this paper cites.
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
X. Li et al · 2024
Earlier work this paper cites.
Vision-language models for vision tasks: A survey
J. Zhang et al · 2024
Earlier work this paper cites.
Mulberry: Empowering mllm with o1-like reasoning and reflection via collective monte carlo tree search, 2024
H. Yao et al · 2024
Earlier work this paper cites.
Openvla: An open-source vision-language-action model
M. J. Kim et al · 2024
Earlier work this paper cites.
Large multimodal agents: A survey
J. Xie et al · 2024
Earlier work this paper cites.
A survey on large language model based autonomous agents
L. Wang et al · 2024
Earlier work this paper cites.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
J. Y. Koh et al · 2024
Earlier work this paper cites.
Adaptagent: Adapting multimodal web agents with few-shot learning from human demonstrations
G. Verma et al · 2024
Earlier work this paper cites.
Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration
J. Wang et al · 2024
Earlier work this paper cites.
A. Dubey et al · 2024
Earlier work this paper cites.
M. Abdin et al · 2024
Earlier work this paper cites.
X. Dong et al · 2024
Earlier work this paper cites.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding
Z. Wu et al · 2024
Earlier work this paper cites.
Mm1. 5: Methods, analysis & insights from multimodal llm fine-tuning
H. Zhang et al · 2024
Earlier work this paper cites.
Vision-language models can self-improve reasoning via reflection
K. Cheng et al · 2024
Earlier work this paper cites.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
B. He et al · 2024
Earlier work this paper cites.
Moviechat: From dense token to sparse memory for long video understanding
E. Song et al · 2024
Earlier work this paper cites.
Longrope: Extending llm context window beyond 2 million tokens
Y. Ding et al · 2024
Earlier work this paper cites.
Longvila: Scaling long-context visual language models for long videos
Y. Chen et al · 2024
Earlier work this paper cites.
Webvoyager: Building an end-to-end web agent with large multimodal models
H. He et al · 2024
Earlier work this paper cites.
Aguvis: Unified pure vision agents for autonomous gui interaction
Y. Xu et al · 2024
Earlier work this paper cites.
Evidential active recognition: Intelligent and prudent open-world embodied perception
L. Fan et al · 2024
Earlier work this paper cites.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Y. Zheng et al · 2024
Earlier work this paper cites.
Swift:a scalable lightweight infrastructure for fine-tuning, 2024
Y. Zhao et al · 2024
Earlier work this paper cites.
Mavis: Mathematical visual instruction tuning
R. Zhang et al · 2024
Earlier work this paper cites.
Gui-world: A video benchmark and dataset for multimodal gui-oriented understanding
D. Chen et al · 2024
Earlier work this paper cites.
Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark
X. Yue et al · 2024
Earlier work this paper cites.
C. He et al · 2024
Earlier work this paper cites.
Mmlongbench-doc: Benchmarking long-context document understanding with visualizations
Y. Ma et al · 2024
Earlier work this paper cites.
Milebench: Benchmarking mllms in long context
D. Song et al · 2024
Earlier work this paper cites.
Design2code: Benchmarking multimodal code generation for automated front-end engineering
C. Si et al · 2024
Earlier work this paper cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
T. Xie et al · 2024
Earlier work this paper cites.
S. Zhang et al · 2024
Earlier work this paper cites.
A. Jaech et al · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao et al · 2024
Earlier work this paper cites.
Autoflow: Automated workflow generation for large language model agents
Z. Li et al · 2024
Earlier work this paper cites.
Videoagent: Long-form video understanding with large language model as agent
X. Wang et al · 2024
Earlier work this paper cites.
Mmsearch: Benchmarking the potential of large models as multi-modal search engines
D. Jiang et al · 2024
Earlier work this paper cites.
Improved baselines with visual instruction tuning
H. Liu et al · 2024
Earlier work this paper cites.
A. Hurst et al · 2024
Earlier work this paper cites.
Mini-gemini: Mining the potential of multi-modality vision language models
Y. Li et al · 2024
Earlier work this paper cites.
Llava-onevision: Easy visual task transfer
B. Li et al · 2024
Earlier work this paper cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
P. Tong et al · 2024
Earlier work this paper cites.
What matters when building vision-language models?
H. Laurençon et al · 2024
Earlier work this paper cites.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
A. Hu et al · 2024
Earlier work this paper cites.
Tinychart: Efficient chart understanding with visual token merging and program-of-thoughts learning
L. Zhang et al · 2024
Earlier work this paper cites.
S. Wang et al · 2024
Earlier work this paper cites.
W. Wang et al · 2024
Earlier work this paper cites.
Automated multi-level preference for mllms
M. Zhang et al · 2024
Earlier work this paper cites.
W. Wang et al · 2024
Earlier work this paper cites.
Moe-llava: Mixture of experts for large vision-language models
B. Lin et al · 2024
Earlier work this paper cites.
Eve: Efficient vision-language pre-training with masked prediction and modality-aware moe
J. Chen et al · 2024
Earlier work this paper cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
D. Dai et al · 2024
Earlier work this paper cites.
Investigating continual pretraining in large language models: Insights and implications
Ç. Yıldız et al · 2024
Earlier work this paper cites.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
R. Zhang et al · 2024
Earlier work this paper cites.
S. Huang et al · 2024
Earlier work this paper cites.
Measuring multimodal mathematical reasoning with math-vision dataset
K. Wang et al · 2024
Earlier work this paper cites.
Androidworld: A dynamic benchmarking environment for autonomous agents
C. Rawles et al · 2024
Earlier work this paper cites.
Dense connector for mllms
H. Yao et al · 2024
Earlier work this paper cites.
Llm maybe longlm: Self-extend llm context window without tuning
H. Jin et al · 2024
Earlier work this paper cites.
Long context transfer from language to vision
P. Zhang et al · 2024
Earlier work this paper cites.
Memory-augmented multimodal llms for surgical vqa via self-contained inquiry
W. Hou et al · 2024
Earlier work this paper cites.
Accelerating best-of-n via speculative rejection
R. Zhang et al · 2024
Earlier work this paper cites.
Forest-of-thought: Scaling test-time compute for enhancing llm reasoning
Z. Bi et al · 2024
Earlier work this paper cites.
Improve vision language model chain-of-thought reasoning
R. Zhang et al · 2024
Earlier work this paper cites.
Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale
J. Guo et al · 2024
Earlier work this paper cites.
Textcot: Zoom in for enhanced multimodal text-rich image understanding
B. Luan et al · 2024
Earlier work this paper cites.
Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
H. Shao et al · 2024
Earlier work this paper cites.
Compositional chain-of-thought prompting for large multimodal models
C. Mitra et al · 2024
Earlier work this paper cites.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
L. Xu et al · 2024
Earlier work this paper cites.
Moviechat+: Question-aware sparse memory for long video question answering
E. Song et al · 2024
Earlier work this paper cites.
Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage
Z. Gao et al · 2024
Earlier work this paper cites.
Vision search assistant: Empower vision-language models as multimodal search engines
Z. Zhang et al · 2024
Earlier work this paper cites.
Mindsearch: Mimicking human minds elicits deep ai searcher
Z. Chen et al · 2024
Earlier work this paper cites.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence, 2024
D. Guo et al · 2024
Earlier work this paper cites.
Stepcoder: Improve code generation with reinforcement learning from compiler feedback
S. Dou et al · 2024
Earlier work this paper cites.
J. Wang et al · 2024
Earlier work this paper cites.
Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
H. Shao et al · 2024
Earlier work this paper cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu et al · 2024
Earlier work this paper cites.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
J. Hu et al · 2024
Earlier work this paper cites.
Hybridflow: A flexible and efficient rlhf framework
G. Sheng et al · 2024
Earlier work this paper cites.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu et al · 2024
Earlier work this paper cites.
We-math: Does your large multimodal model achieve human-like mathematical reasoning?
R. Qiao et al · 2024
Earlier work this paper cites.
Charxiv: Charting gaps in realistic chart understanding in multimodal llms
Z. Wang et al · 2024
Earlier work this paper cites.
Evaluating very long-term conversational memory of llm agents
A. Maharana et al · 2024
Earlier work this paper cites.
Longvideobench: A benchmark for long-context interleaved video-language understanding
H. Wu et al · 2024
Cited alongside, same era.
Lvbench: An extreme long video understanding benchmark
W. Wang et al · 2024
Cited alongside, same era.
V?: Guided visual search as a core mechanism in multimodal llms
P. Wu and S. Xie · 2024
Cited alongside, same era.
Seeclick: Harnessing gui grounding for advanced visual gui agents
K. Cheng et al · 2024
Cited alongside, same era.
On the effects of data scale on ui control agents
W. Li et al · 2024
Cited alongside, same era.
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web
Mobilerl: Online agentic reinforcement learning for mobile gui agents, 2025
Y. Xu et al · 2025
Closest in time.
L. Lin et al · 2025
Closest in time.
Ui-r1: Enhancing action prediction of gui agents by reinforcement learning
Z. Lu et al · 2025
Closest in time.
Drive-r1: Bridging reasoning and planning in vlms for autonomous driving with reinforcement learning
Y. Li et al · 2025
Closest in time.
B. Jiang et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Kapoor et al · 2024
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black et al · 2024
Cited alongside, same era.
Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale
J. Chen et al · 2024
Cited alongside, same era.
Huatuogpt-o1, towards medical complex reasoning with llms
J. Chen et al · 2024
Cited alongside, same era.
Mmed-rag: Versatile multimodal rag system for medical vision language models
P. Xia et al · 2024
Cited alongside, same era.
A fully autonomous robotic ultrasound system for thyroid scanning
K. Su et al · 2024
Cited alongside, same era.
Gui odyssey: A comprehensive dataset for cross-app gui navigation on mobile devices
Q. Lu et al · 2024
Cited alongside, same era.
Closest in time.
K. Qian et al · 2025
Closest in time.
X. Hou et al · 2025
Closest in time.
Towards agentic recommender systems in the era of multimodal large language models
C. Huang et al · 2025
Closest in time.
Recoworld: Building simulated environments for agentic recommender systems
F. Liu et al · 2025
Closest in time.
Vragent-r1: Boosting video recommendation with mllm-based agents via reinforcement learning
S. Chen et al · 2025
Closest in time.
Reasonrec: A reasoning-augmented multimodal agent for unified recommendation
Y. Zhang et al · 2025
Closest in time.
A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems
Z. Ke et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo et al · 2025
Closest in time.
The landscape of agentic reinforcement learning for llms: A survey
G. Zhang et al · 2025
Closest in time.
A survey of reinforcement learning for large reasoning models
K. Zhang et al · 2025
Closest in time.
Miroflow: An open-source agentic framework for deep research
MiroMind AI Team · 2025
Closest in time.
Introducing perplexity deep research
Perplexity Team · 2025
Closest in time.
Agentgroupchat-v2: Divide-and-conquer is what llm-based multi-agent system need
Z. Gu et al · 2025
Closest in time.
Ocean-ocr: Towards general ocr application via a vision-language model
S. Chen et al · 2025
Closest in time.
Y. Feng et al · 2025
Closest in time.
Valley2: Exploring multimodal models with scalable vision-language design
Z. Wu et al · 2025
Closest in time.
A. Yang et al · 2025
Closest in time.
Glm-4.5: Agentic, reasoning, and coding (arc) foundation models
A. Zeng et al · 2025
Closest in time.
Finevision: Open data is all you need, September 2025
L. Wiedmann et al · 2025
Closest in time.
Be confident: Uncovering overfitting in mllm multi-task tuning
W. Huang et al · 2025
Closest in time.
Uni-moe: Scaling unified multimodal llms with mixture of experts
Y. Li et al · 2025
Closest in time.
Cl-moe: Enhancing multimodal large language model with dual momentum mixture-of-experts for continual visual question answering
T. Huai et al · 2025
Closest in time.
gpt-oss-120b & gpt-oss-20b model card
S. Agarwal et al · 2025
Closest in time.
M. Wang et al · 2025
Closest in time.
Scaling agents via continual pre-training
L. Su et al · 2025
Closest in time.
Comp: Continual multimodal pre-training for vision foundation models
Y. Chen et al · 2025
Closest in time.
Masksearch: A universal pre-training framework to enhance agentic search capability
W. Wu et al · 2025
Closest in time.
Webwalker: Benchmarking llms in web traversal
J. Wu et al · 2025
Closest in time.
K. Li et al · 2025
Closest in time.
Websailor: Navigating super-human reasoning for web agent
K. Li et al · 2025
Closest in time.
Webshaper: Agentically data synthesizing via information-seeking formalization
Z. Tao et al · 2025
Closest in time.
D. Jiang et al · 2025
Closest in time.
Mmreason: An open-ended multi-modal multi-step reasoning benchmark for mllms toward agi
H. Yao et al · 2025
Closest in time.
Echoink-r1: Exploring audio-visual reasoning in multimodal llms via reinforcement learning
Z. Xing et al · 2025
Closest in time.
Z. Liu et al · 2025
Closest in time.
Noisyrollout: Reinforcing visual reasoning with data augmentation
X. Liu et al · 2025
Closest in time.
T. Xiao et al · 2025
Closest in time.
Visualprm: An effective process reward model for multimodal reasoning
W. Wang et al · 2025
Closest in time.
Mm-prm: Enhancing multimodal mathematical reasoning with scalable step-level supervision
L. Du et al · 2025
Closest in time.
Rm-r1: Reward modeling as reasoning
X. Chen et al · 2025
Closest in time.
Scalable best-of-n selection for large language models via self-certainty
Z. Kang et al · 2025
Closest in time.
Visuothink: Empowering lvlm reasoning with multimodal tree search, 2025
Y. Wang et al · 2025
Closest in time.
Boosting multimodal reasoning with automated structured thinking
J. Wu et al · 2025
Closest in time.
Corvid: Improving multimodal large language models towards chain-of-thought reasoning
J. Jiang et al · 2025
Closest in time.
Insight-v: Exploring long-chain visual reasoning with multimodal large language models
Y. Dong et al · 2025
Closest in time.
Llamav-o1: Rethinking step-by-step visual reasoning in llms, 2025
O. Thawakar et al · 2025
Closest in time.
R1-compress: Long chain-of-thought compression via chunk compression and search
Y. Wang et al · 2025
Closest in time.
Inference-time reward hacking in large language models
H. Khalaf et al · 2025
Closest in time.
Beyond reward hacking: Causal rewards for large language model alignment
C. Wang et al · 2025
Closest in time.
Seg-zero: Reasoning-chain guided segmentation via cognitive reinforcement
Y. Liu et al · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Y. Peng et al · 2025
Closest in time.
Y. Zhan et al · 2025
Closest in time.
Tiny-r1v: Lightweight multimodal unified reasoning model via model merging, 2025
Q. Yin et al · 2025
Closest in time.
Mapo: Mixed advantage policy optimization
W. Huang et al · 2025
Closest in time.
H. Deng et al · 2025
Closest in time.
Posterior-grpo: Rewarding reasoning processes in code generation
L. Fan et al · 2025
Closest in time.
R1-code-interpreter: Training llms to reason with code via supervised and reinforcement learning
Y. Chen et al · 2025
Closest in time.
Otc: Optimal tool calls via reinforcement learning
H. Wang et al · 2025
Closest in time.
rstar2-agent: Agentic reasoning technical report
N. Shang et al · 2025
Closest in time.
Medagentgym: Training llm agents for code-based medical reasoning at scale
R. Xu et al · 2025
Closest in time.
Ml-agent: Reinforcing llm agents for autonomous machine learning engineering
Z. Liu et al · 2025
Closest in time.
Reinforcing spatial reasoning in vision-language models with interwoven thinking and visual drawing
J. Wu et al · 2025
Closest in time.
Mllm-tool: A multimodal large language model for tool agent learning
C. Wang et al · 2025
Closest in time.
Introducing gpt-5
OpenAI · 2025
Closest in time.
Ui-tars-2 technical report: Advancing gui agent with multi-turn reinforcement learning
H. Wang et al · 2025
Closest in time.
Acereason-nemotron: Advancing math and code reasoning through reinforcement learning
Y. Chen et al · 2025
Closest in time.
Towards better correctness and efficiency in code generation
Y. Feng et al · 2025
Closest in time.
Process-supervised reinforcement learning for code generation
Y. Ye et al · 2025
Closest in time.
Co-evolving llm coder and unit tester via reinforcement learning
Y. Wang et al · 2025
Closest in time.
Os-r1: Agentic operating system kernel tuning with reinforcement learning
H. Lin et al · 2025
Closest in time.
Thinking with images
OpenAI · 2025
Closest in time.
Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning
C. Duan et al · 2025
Closest in time.
Can we generate images with cot? let’s verify and reinforce image generation step by step
Z. Guo et al · 2025
Closest in time.
Delving into rl for image generation with cot: A study on dpo vs. grpo
C. Tong et al · 2025
Closest in time.
Open r1: A fully open reproduction of deepseek-r1, January 2025
Hugging Face · 2025
Closest in time.
Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025
T. Xie et al · 2025
Closest in time.
Easyr1: An efficient, scalable, multi-modality rl training framework
Y. Zheng et al · 2025
Closest in time.
7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
W. Zeng et al · 2025
Closest in time.
Light-r1: Curriculum sft, dpo and rl for long cot from scratch and beyond
L. Wen et al · 2025
Closest in time.
Areal: A large-scale asynchronous reinforcement learning system for language reasoning, 2025
W. Fu et al · 2025
Closest in time.
Visual agentic reinforcement fine-tuning, 2025
Z. Liu et al · 2025
Closest in time.
Search-r1: Training llms to reason and leverage search engines with reinforcement learning
B. Jin et al · 2025
Closest in time.
Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025
Z. Wang et al · 2025
Closest in time.
Marti: A framework for multi-agent llm systems reinforced training and inference, 2025
K. Zhang et al · 2025
Closest in time.
Mirorl: An mcp-first reinforcement learning framework for deep research agent
M. F. M. Team and M. A. I. Team · 2025
Closest in time.
Aworld: Orchestrating the training recipe for agentic ai, 2025
C. Yu et al · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Y. Yang et al · 2025
Closest in time.
Advancing multimodal reasoning: From optimized cold start to staged reinforcement learning
S. Chen et al · 2025
Closest in time.
Video-xl-pro: Reconstructive token compression for extremely long video understanding
X. Liu et al · 2025
Closest in time.
R1-searcher: Incentivizing the search capability in llms via reinforcement learning
H. Song et al · 2025
Closest in time.
rstar-coder: Scaling competitive code reasoning with a large-scale verified dataset
Y. Liu et al · 2025
Closest in time.
Chain-of-focus: Adaptive visual search and zooming for multimodal reasoning via rl
X. Zhang et al · 2025
Closest in time.
Instructvla: Vision-language-action instruction tuning from understanding to manipulation
S. Yang et al · 2025
Closest in time.
Webdancer: Towards autonomous information seeking agency
J. Wu et al · 2025
Closest in time.
Zerobench: An impossible visual benchmark for contemporary large multimodal models
J. Roberts et al · 2025
Closest in time.
Zebralogic: On the scaling limits of llms for logical reasoning
B. Y. Lin et al · 2025
Closest in time.
Videomathqa: Benchmarking mathematical reasoning via multimodal understanding in videos
H. Rasheed et al · 2025
Closest in time.
L. Phan et al · 2025
Closest in time.
Vidorag: Visual document retrieval-augmented generation via dynamic iterative reasoning agents
Q. Wang et al · 2025
Closest in time.
Advancing vision-language models in front-end development via data synthesis
T. Ge et al · 2025
Closest in time.
Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
W. Wang et al · 2025
Closest in time.
Screenspot-pro: Gui grounding for professional high-resolution computer use
K. Li et al · 2025
Closest in time.
Y. Dong et al · 2025
Closest in time.
A comprehensive survey of deep research: Systems, methodologies, and applications, 2025
R. Xu and J. Peng · 2025
Closest in time.
Deep research agents: A systematic examination and roadmap
Y. Huang et al · 2025
Closest in time.
From web search towards agentic deep research: Incentivizing search with reasoning agents
W. Zhang et al · 2025
Closest in time.
Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025
Y. Zheng et al · 2025
Closest in time.
Introducing researcher and analyst in microsoft 365 copilot
Microsoft · 2025
Closest in time.
Kimi-researcher: End-to-end rl training for emerging agentic capabilities
Moonshot AI · 2025
Closest in time.
Autoglm rumination
Zhipu AI · 2025
Closest in time.
Manus: General ai agent that bridges mind and action
Manus Team · 2025
Closest in time.
Univla: Learning to act anywhere with task-centric latent actions
Q. Bu et al · 2025
Closest in time.
F1: A vision-language-action model bridging understanding and generation to actions
Q. Lv et al · 2025
Closest in time.
C. Cheang et al · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
K. Black et al · 2025
Closest in time.
Molmoact: Action reasoning models that can reason in space
J. Lee et al · 2025
Closest in time.
Embodied ai: Emerging risks and opportunities for policy action
J. Perlo et al · 2025
Closest in time.
Aligning cyber space with physical world: A comprehensive survey on embodied ai
Y. Liu et al · 2025
Closest in time.
A. Yu et al · 2025
Closest in time.
W. Zhang et al · 2025
Closest in time.
At-cxr: Uncertainty-aware agentic triage for chest x-rays
X. Li et al · 2025
Closest in time.
Gem: Empowering mllm for grounded ecg understanding with time series and images
X. Lan et al · 2025
Closest in time.
Mobile-agent-v3: Foundamental agents for gui automation
J. Ye et al · 2025
Closest in time.
Ui-vision: A desktop-centric gui benchmark for visual perception and interaction
S. Nayak et al · 2025
Closest in time.
Worldgui: An interactive benchmark for desktop gui automation from any starting point
H. H. Zhao et al · 2025
Closest in time.
Fingertip 20k: A benchmark for proactive and personalized mobile llm agents
Q. Yang et al · 2025
Closest in time.
Chain-of-thought for autonomous driving: A comprehensive survey and future prospects
Y. Cui et al · 2025
Closest in time.
Z. Yuan et al · 2025
Closest in time.
A. Ishaq et al · 2025
Closest in time.
Towards agentic recommender systems in the era of multimodal large language models
C. Huang et al · 2025
Closest in time.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
H. Luo et al · 2025
Closest in time.
Ada-r1: Hybrid-cot via bi-level adaptive reasoning optimization
H. Luo et al · 2025
Closest in time.
Fast-slow thinking for large vision-language model reasoning
W. Xiao et al · 2025
Closest in time.
Adacot: Pareto-optimal adaptive chain-of-thought triggering via reinforcement learning
C. Lou et al · 2025
Closest in time.
Mla-trust: Benchmarking trustworthiness of multimodal llm agents in gui environments
X. Yang et al · 2025
Closest in time.
S. Raza et al · 2025
Closest in time.