Fetching the paper…
Reading the bibliography…
Retrieval-augmented generation (RAG) has emerged as a pivotal technique in artificial intelligence (AI), particularly in enhancing the capabilities of large language models (LLMs) by enabling access to external, reliable, and up-to-date knowledge sources.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision – ECCV 2014 , D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds. Cham: Springer International Publishing, 2014, pp. 740–755
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13 . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 . IEEE Computer Society, 2017, pp. 6325–6334
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in CVPR . IEEE Computer Society, 2017, pp. 6325–6334
2017
Earlier work this paper cites.
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018 . IEEE Computer Society, 2018, pp. 3608–3617
2018
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019, pp. 2630–2640
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 . Computer Vision Foundation / IEEE, 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Gupta, P. Dollár, and R. B. Girshick, “LVIS: A dataset for large vocabulary instance segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 . Computer Vision Foundation / IEEE, 2019, pp. 5356–5364
2019
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” NeurIPS , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
N. Spolaõr, H. D. Lee, W. S. R. Takaki, L. A. Ensina, C. S. R. Coy, and F. C. Wu, “A systematic review on content-based video retrieval,” in Eng. Appl. Artif. Intell. , 2020
2020
Earlier work this paper cites.
M. Mathew, D. Karatzas, and C. V. Jawahar, “Docvqa: A dataset for vqa on document images,” in WACV , 2021, pp. 2199–2208
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Dhamo, F. Manhardt, N. Navab, and F. Tombari, “Graph-to-3d: End-to-end generation and manipulation of 3d scenes using scene graphs,” in ICCV , 2021, pp. 16 352–16 361
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Long, W. Yin, T. Ajanthan, V. Nguyen, P. Purkait, R. Garg, A. Blair, C. Shen, and A. van den Hengel, “Retrieval augmented classification for long-tail visual recognition,” in CVPR , 2022, pp. 6959–6969
2022
Earlier work this paper cites.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in Computer Vision – ECCV 2022 , S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds. Cham: Springer Nature Switzerland, 2022, pp. 146–162
2022
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K. Chang, S. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in NeurIPS 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., 2022
2022
Earlier work this paper cites.
C. Li, H. Liu, L. H. Li, P. Zhang, J. Aneja, J. Yang, P. Jin, H. Hu, Z. Liu, Y. J. Lee, and J. Gao, “ELEVATER: A benchmark and toolkit for evaluating language-augmented visual models,” in NeurIPS 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., 2022
2022
Earlier work this paper cites.
W. Chen, H. Hu, X. Chen, P. Verga, and W. Cohen, “MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,” in EMNLP , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: ACL, 2022, pp. 5558–5570
2022
Earlier work this paper cites.
A. Blattmann, R. Rombach, K. Oktay, J. Müller, and B. Ommer, “Retrieval-augmented diffusion models,” NeurIPS , vol. 35, pp. 15 309–15 324, 2022
2022
Earlier work this paper cites.
Y. Chang, M. Narang, H. Suzuki, G. Cao, J. Gao, and Y. Bisk, “Webqa: Multihop and multimodal qa,” in CVPR , 2022, pp. 16 495–16 504
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS , vol. 35, pp. 25 278–25 294, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Piadyk, J. Rulff, E. Brewer, M. Hosseini, K. Ozbay, M. Sankaradas, S. Chakradhar, and C. Silva, “Streetaware: A high-resolution synchronized multimodal urban scene dataset,” Sensors , vol. 23, no. 7, p. 3710, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Zhou and G. Long, “Style-aware contrastive learning for multi-style image captioning,” in EACL , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” in arxiv , 2023
2023
Earlier work this paper cites.
Y. He, M. Xia, H. Chen, X. Cun, Y. Gong, J. Xing, Y. Zhang, X. Wang, C. Weng, Y. Shan et al. , “Animate-a-story: Storytelling with retrieval-augmented video generation,” arXiv , 2023
2023
Earlier work this paper cites.
H. Hu, Y. Luan, Y. Chen, U. Khandelwal, M. Joshi, K. Lee, K. Toutanova, and M. Chang, “Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities,” in IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 . IEEE, 2023, pp. 12 031–12 041
2023
Earlier work this paper cites.
Y. Chen, H. Hu, Y. Luan, H. Sun, S. Changpinyo, A. Ritter, and M.-W. Chang, “Can pre-trained vision and language models answer visual information-seeking questions?” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: ACL, 2023, pp. 14 948–14 968
2023
Earlier work this paper cites.
J. Wang, P. Zhang, T. Chu, Y. Cao, Y. Zhou, T. Wu, B. Wang, C. He, and D. Lin, “V3det: Vast vocabulary visual detection dataset,” in IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 . IEEE, 2023, pp. 19 787–19 797
2023
Earlier work this paper cites.
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin, “Mmbench: Is your multi-modal model an all-around player?” 2023
2023
Earlier work this paper cites.
Y. Li, Y. Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: ACL, 2023, pp. 292–305
2023
Earlier work this paper cites.
H. Liu, K. Son, J. Yang, C. Liu, J. Gao, Y. J. Lee, and C. Li, “Learning customized visual models with retrieval-augmented knowledge,” in CVPR . IEEE, 2023, pp. 15 148–15 158
2023
Earlier work this paper cites.
W. Lin, J. Chen, J. Mei, A. Coca, and B. Byrne, “Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering,” in NeurIPS , 2023
2023
Earlier work this paper cites.
J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, Y. Dan, C. Zhao, G. Xu, C. Li, J. Tian, Q. Qi, J. Zhang, and F. Huang, “mplug-docowl: Modularized multimodal large language model for document understanding,” 2023
2023
Earlier work this paper cites.
Z. Hu, A. Iscen, C. Sun, Z. Wang, K. Chang, Y. Sun, C. Schmid, D. A. Ross, and A. Fathi, “Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory,” in CVPR . IEEE, 2023, pp. 23 369–23 379
2023
Earlier work this paper cites.
J. Rao, Z. Shan, L. Liu, Y. Zhou, and Y. Yang, “Retrieval-based knowledge augmented vision language pre-training,” in ACM MM . ACM, 2023, pp. 5399–5409
2023
Earlier work this paper cites.
D. Cioni, L. Berlincioni, F. Becattini, and A. Del Bimbo, “Diffusion based augmentation for captioning and retrieval in cultural heritage,” in ICCV , 2023, pp. 1707–1716
2023
Earlier work this paper cites.
M. Zhang, X. Guo, L. Pan, Z. Cai, F. Hong, H. Li, L. Yang, and Z. Liu, “Remodiffuse: Retrieval-augmented motion diffusion model,” in ICCV , 2023, pp. 364–373
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in CVPR , 2023, pp. 13 142–13 153
2023
Cited alongside, same era.
M. Zhang, X. Guo, L. Pan, Z. Cai, F. Hong, H. Li, L. Yang, and Z. Liu, “Remodiffuse: Retrieval-augmented motion diffusion model,” in ICCV (ICCV) , October 2023, pp. 364–373
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Fan, Y. Ding, L. Ning, S. Wang, H. Li, D. Yin, T. Chua, and Q. Li, “A survey on RAG meeting llms: Towards retrieval-augmented large language models,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2024, Barcelona, Spain, August 25-29, 2024 , R. Baeza-Yates and F. Bonchi, Eds. ACM, 2024, pp. 6491–6501
Z. Wang, M. Xia, L. He, H. Chen, Y. Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, A. Chevalier, S. Arora, and D. Chen, “Charxiv: Charting gaps in realistic chart understanding in multimodal LLMs,” in NeurIPS , 2024
2024
Later among the works it cites.
V. N. Rao, S. Choudhary, A. Deshpande, R. K. Satzoda, and S. Appalaraju, “Raven: Multitask retrieval augmented vision-language learning,” 2024
2024
Later among the works it cites.
J. Sun, J. Zhang, Y. Zhou, Z. Su, X. Qu, and Y. Cheng, “SURf: Teaching large vision-language models to selectively utilize retrieved information,” in EMNLP , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: ACL, 2024, pp. 7611–7629
2024
Later among the works it cites.
J. Qi, Z. Xu, R. Shao, Y. Chen, J. Di, Y. Cheng, Q. Wang, and L. Huang, “Rora-vlm: Robust retrieval-augmented vision language models,” 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
H. Yu, A. Gan, K. Zhang, S. Tong, Q. Liu, and Z. Liu, “Evaluation of retrieval-augmented generation: A survey,” in CCF Conference on Big Data . Springer, 2024, pp. 102–120
2024
Cited alongside, same era.
T. T. Procko and O. Ochoa, “Graph retrieval-augmented generation for large language models: A survey,” in 2024 Conference on AI, Science, Engineering, and Technology (AIxSET) . IEEE, 2024, pp. 166–169
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
P. Xia, K. Zhu, H. Li, H. Zhu, Y. Li, G. Li, L. Zhang, and H. Yao, “Rule: Reliable multimodal rag for factuality in medical vision language models,” in EMNLP , 2024, pp. 1081–1093
2024
Later among the works it cites.
“Image-based rag: https://baike.baidu.com/item/iRAG/65102065?fr=ge_ala ,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Horita, N. Inoue, K. Kikuchi, K. Yamaguchi, and K. Aizawa, “Retrieval-augmented layout transformer for content-aware layout generation,” in CVPR , 2024, pp. 67–76
2024
Later among the works it cites.
J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “Lgm: Large multi-view gaussian model for high-resolution 3d content creation,” in ECCV . Springer, 2024, pp. 1–18
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V. Voleti, S. Y. Gadre et al. , “Objaverse-xl: A universe of 10m+ 3d objects,” NeurIPS , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Q. Wu, D. Iliash, D. Ritchie, M. Savva, and A. X. Chang, “Diorama: Unleashing zero-shot single-view 3d scene modeling,” 2024
2024
Later among the works it cites.
W. Xu, M. Wang, W. Zhou, and H. Li, “P-rag: Progressive retrieval augmented generation for planning on embodied everyday task,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 6969–6978
2024
Later among the works it cites.
Y. Zhu, Z. Ou, X. Mou, and J. Tang, “Retrieval-augmented embodied agents,” in CVPR , 2024, pp. 17 985–17 995
2024
Later among the works it cites.
2024
Later among the works it cites.
Q. Xie, S. Y. Min, T. Zhang, K. Xu, A. Bajaj, R. Salakhutdinov, M. Johnson-Roberson, and Y. Bisk, “Embodied-rag: General non-parametric embodied memory for retrieval and generation,” in Language Gamification-NeurIPS 2024 Workshop , 2024
2024
Later among the works it cites.
W. Ding, Y. Cao, D. Zhao, C. Xiao, and M. Pavone, “Realgen: Retrieval augmented generation for controllable traffic scenarios,” in ECCV . Springer, 2024, pp. 93–110
2024
Later among the works it cites.
2024
Later among the works it cites.
“Retrieval augmented generation: 40% reduction in inspection times and notable improvement in defect detection rates,” 2024
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
S. Jeong, K. Kim, J. Baek, and S. J. Hwang, “Videorag: Retrieval-augmented generation over video corpus,” 2025
2025
Closest in time.
M. Sankaradas, R. K. Rajendran, and S. T. Chakradhar, “Streamingrag: Real-time contextual retrieval and generation framework,” arXiv, 2025
2025
Closest in time.
M. Bonomo and S. Bianco, “Visual rag: Expanding mllm visual knowledge without fine-tuning,” 2025
2025
Closest in time.
H. Xiong, Z. Yang, J. Yu, Y. Zhuge, L. Zhang, J. Zhu, and H. Lu, “Streaming video understanding and multi-round interaction with memory-enhanced knowledge,” 2025
2025
Closest in time.
H. Fang, X. Sui, H. Yu, J. Kong, S. Yu, B. Chen, H. Wu, and S.-T. Xia, “Retrievals can be detrimental: A contrastive backdoor attack paradigm on retrieval-augmented diffusion models,” 2025
2025
Closest in time.
B. Luo, J. Wang, Z. Wang, J. Zhu, and X. Zhao, “Graph-based cross-domain knowledge distillation for cross-dataset text-to-image person retrieval,” arXiv, 2025
2025
Closest in time.
Y. Xu, Y. Sun, B. Zhai, M. Li, W. Liang, Y. Li, and S. Du, “Zero-shot video moment retrieval via off-the-shelf multimodal large language models,” 2025
2025
Closest in time.
M. Faysse, H. Sibille, T. Wu, B. Omrani, G. Viaud, C. HUDELOT, and P. Colombo, “Colpali: Efficient document retrieval with vision language models,” in ICLR , 2025
2025
Closest in time.
K. Dong, Y. Chang, X. D. Goh, D. Li, R. Tang, and Y. Liu, “Mmdocir: Benchmarking multi-modal retrieval for long documents,” 2025
2025
Closest in time.
Y. Wu, Q. Long, J. Li, J. Yu, and W. Wang, “Visual-rag: Benchmarking text-to-image retrieval augmented generation for visual knowledge intensive queries,” 2025
2025
Closest in time.
W. Hu, J.-C. Gu, Z.-Y. Dou, M. Fayyaz, P. Lu, K.-W. Chang, and N. Peng, “MRAG-bench: Vision-centric evaluation for retrieval-augmented multimodal models,” in ICLR , 2025
2025
Closest in time.
N. Wasserman, R. Pony, O. Naparstek, A. R. Goldfarb, E. Schwartz, U. Barzelay, and L. Karlinsky, “Real-mm-rag: A real-world multi-modal retrieval benchmark,” 2025
2025
Closest in time.
M. Mortaheb, M. A. A. Khojastepour, S. T. Chakradhar, and S. Ulukus, “Rag-check: Evaluating multimodal retrieval augmented generation performance,” arXiv , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Zhang, Z. Chong, X. Zhang, H. Li, Y. Cheng, Y. Yan, and X. Liang, “Garmentaligner: Text-to-garment generation via retrieval-augmented multi-level corrections,” in ECCV . Springer, 2025, pp. 148–164
2025
Closest in time.
W. Ding, Y. Cao, D. Zhao, C. Xiao, and M. Pavone, “Realgen: Retrieval augmented generation for controllable traffic scenarios,” in ECCV . Springer, 2025, pp. 93–110
2025
Closest in time.
H. Yuan, Z. Zhao, S. Wang, S. Xiao, M. Ni, Z. Liu, and Z. Dou, “Finerag: Fine-grained retrieval-augmented text-to-image generation,” in Proceedings of the 31st International Conference on Computational Linguistics , 2025, pp. 11 196–11 205
2025
Closest in time.