Fetching the paper…
Reading the bibliography…
Current research on Multimodal Retrieval-Augmented Generation (MRAG) enables diverse multimodal inputs but remains limited to single-modality outputs, restricting expressive capacity and practical utility.
Bertscore: Evaluating text generation with bert
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
ELI5: Long form question answering
Fan, A.; Jernite, Y.; Perez, E.; Grangier, D.; Weston, J.; and Auli, M. 2019 · 1907
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; et al. 2020 · 2020
Earlier work this paper cites.
Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models
Bao, F.; Li, C.; Zhu, J.; and Zhang, B. 2022 · 2022
Earlier work this paper cites.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
Chen, W.; Hu, H.; Chen, X.; Verga, P.; and Cohen, W. W. 2022 · 2022
Earlier work this paper cites.
Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; et al. 2023 · 2023
Earlier work this paper cites.
Learning customized visual models with retrieval-augmented knowledge
Liu, H.; Son, K.; Yang, J.; Liu, C.; Gao, J.; Lee, Y. J.; and Li, C. 2023 · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; Millican, K.; et al. 2023 · 2023
Earlier work this paper cites.
Claude 3.5 Sonnet
Anthropic. 2024 · 2024
Earlier work this paper cites.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; and Hajishirzi, H. 2024 · 2024
Earlier work this paper cites.
AI-thenticity: Exploring the effect of perceived authenticity of AI-generated visual content on tourist patronage intentions
Bui, H. T.; Filimonau, V.; and Sezerel, H. 2024 · 2024
Earlier work this paper cites.
Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. 2024 · 2024
Cited alongside, same era.
Autoregressive image generation without vector quantization
Li, T.; Tian, Y.; Li, H.; Deng, M.; and He, K. 2024 · 2024
Cited alongside, same era.
Ma, Z.-A.; Lan, T.; Tu, R.-C.; Hu, Y.; Huang, H.; and Mao, X.-L. 2024 · 2024
Cited alongside, same era.
Hello GPT-4o
OpenAI. 2024 · 2024
Cited alongside, same era.
Ftii-bench: A comprehensive multimodal benchmark for flow text with image insertion
Ruan, J.; Yang, Y.; Lin, Z.; Feng, Y.; Xiong, F.; Tang, Z.; and Li, Z. 2024 · 2024
Cited alongside, same era.
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Zhou, J.; Liu, Z.; Liu, Z.; Xiao, S.; Wang, Y.; Zhao, B.; Zhang, C. J.; Lian, D.; and Xiong, Y. 2024 · 2024
Later among the works it cites.
Zhu, Z.; Lee, D.; Zhang, H.; Harsha, S. S.; Feujio, L.; Maharaj, A.; and Li, Y. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025 · 2025
Closest in time.
Vision-r1: Incentivizing reasoning capability in multimodal large language models
Huang, W.; Jia, B.; Zhai, Z.; Cao, S.; Ye, Z.; Zhao, F.; Xu, Z.; Hu, Y.; and Lin, S. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. 2024 · 2024
Cited alongside, same era.
HybridFlow: A Flexible and Efficient RLHF Framework
Sheng, G.; Zhang, C.; Ye, Z.; Wu, X.; Zhang, W.; Zhang, R.; Peng, Y.; Lin, H.; and Wu, C. 2024 · 2024
Cited alongside, same era.
Generative multimodal models are in-context learners
Sun, Q.; Cui, Y.; Zhang, X.; Zhang, F.; Yu, Q.; Wang, Y.; Rao, Y.; Liu, J.; Huang, T.; and Wang, X. 2024 · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G.; Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. 2024 · 2024
Cited alongside, same era.
A survey on multimodal large language models
Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; and Chen, E. 2024 · 2024
Cited alongside, same era.
Visrag: Vision-based retrieval-augmented generation on multi-modality documents
Yu, S.; Tang, C.; Xu, B.; Cui, J.; Ran, J.; Yan, Y.; Liu, Z.; Wang, S.; Han, X.; Liu, Z.; et al. 2024 · 2024
Cited alongside, same era.
Retrieval-augmented generation for ai-generated content: A survey
Zhao, P.; Zhang, H.; Yu, Q.; Wang, Z.; Geng, Y.; Fu, F.; Yang, L.; Zhang, W.; and Cui, B. 2024 · 2024
Cited alongside, same era.
Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; and Han, J. 2025 · 2025
Closest in time.
Visual Agentic Reinforcement Fine-Tuning
Liu, Z.; Zang, Y.; Zou, Y.; Liang, Z.; Dong, X.; Cao, Y.; Duan, H.; Lin, D.; and Wang, J. 2025 · 2025
Closest in time.
A survey of multimodal retrieval-augmented generation
Mei, L.; Mo, S.; Yang, Z.; and Chen, C. 2025 · 2025
Closest in time.
Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning
Meng, F.; Du, L.; Liu, Z.; Zhou, Z.; Lu, Q.; Fu, D.; Shi, B.; Wang, W.; He, J.; Zhang, K.; et al. 2025 · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Peng, Y.; Zhang, G.; Zhang, M.; You, Z.; Liu, J.; Zhu, Q.; Yang, K.; Xu, X.; Geng, X.; and Yang, X. 2025 · 2025
Closest in time.
Wang, Q.; Ding, R.; Zeng, Y.; Chen, Z.; Chen, L.; Wang, S.; Xie, P.; Huang, F.; and Zhao, F. 2025 · 2025
Closest in time.
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation
Yu, Q.; Xiao, Z.; Li, B.; Wang, Z.; Chen, C.; and Zhang, W. 2025 · 2025
Closest in time.
Easyr1: An efficient, scalable, multi-modality rl training framework
Zheng, Y.; Lu, J.; Wang, S.; Feng, Z.; Kuang, D.; and Xiong, Y. 2025 · 2025
Closest in time.
Reinforced mllm: A survey on rl-based reasoning in multimodal large language models
Zhou, G.; Qiu, P.; Chen, C.; Wang, J.; Yang, Z.; Xu, J.; and Qiu, M. 2025 · 2025
Closest in time.