Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs
Alistair EW Johnson, Tom J Pollard, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Yifan Peng, et al. 2019 · 1901
Earlier work this paper cites.
Deep learning for image super-resolution: A survey
Zhihao Wang, Jian Chen, and Steven C. H. Hoi. 2020b · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 1904
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Fashionpedia: Ontology, segmentation, and an attribute localization dataset
Menglin Jia, Mengyun Shi, Mikhail Sirotenko, Yin Cui, Claire Cardie, Bharath Hariharan, Hartwig Adam, and Serge Belongie. 2020 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Maxsim: A maximum similarity metric for machine translation evaluation
Yee Seng Chan and Hwee Tou Ng. 2008 · 2008
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L Chen and William B Dolan. 2011 · 2011
Earlier work this paper cites.
On modality bias in the tvqa dataset
Thomas Winterbottom, Sarah Xiao, Alistair McLean, and Noura Al Moubayed. 2020 · 2012
Earlier work this paper cites.
Bonferroni Correction , pages 154–154
Winston Haynes. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, et al. 2015 · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Nihar Tandon, and Bernt Schiele. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, et al. 2016 · 2016
Earlier work this paper cites.
Deepfashion: Powering robust clothes recognition and retrieval with rich annotations
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Audioset: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, et al. 2017 · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Demystifying MMD GANs
Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton. 2018 · 2018
Earlier work this paper cites.
Soccernet: A scalable dataset for action spotting in soccer videos
Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. 2018 · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. 2018 · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata. 2018 · 2018
Earlier work this paper cites.
Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, and Hung-yi Lee. 2018 · 2018
Earlier work this paper cites.
Fashion-gen: The generative fashion dataset and challenge
Negar Rostamzadeh, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Charadesego: A dataset for egocentric video understanding
Gunnar A Sigurdsson, Gul Varol, Giovanni Maria Farinella, et al. 2018 · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso. 2018 · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, et al. 2019 · 2019
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019 · 2019
Earlier work this paper cites.
Eli5: Long form question answering
Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes. 2019 · 2019
Earlier work this paper cites.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019 · 2019
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2019 · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Earlier work this paper cites.
Coco-cn for cross-lingual image tagging, captioning and retrieval
Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, and Jieping Xu. 2019 · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Xinlei Chen, Abhinav Gupta, Marcus Rohrbach, and Devi Parikh. 2019 · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Earlier work this paper cites.
A survey on biomedical image captioning
John Pavlopoulos, Vasiliki Kougia, and Ion Androutsopoulos. 2019 · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Xiaohan Wang, Jingdong Wang, et al. 2019 · 2019
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019 · 2019
Earlier work this paper cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, et al. 2019 · 2019
Earlier work this paper cites.
Cross-modal self-attention network for referring image segmentation
Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. 2019 · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. 2019 · 2019
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020 · 2020
Earlier work this paper cites.
Accelerating large-scale inference with anisotropic vector quantization
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020 · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020 · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gul Varol, and Andrew Zisserman. 2021 · 2021
Earlier work this paper cites.
Viton-hd: High-resolution virtual try-on via image translation
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. 2021 · 2021
Earlier work this paper cites.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Earlier work this paper cites.
Paq: 65 million probably-asked questions and what you can do with them
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Earlier work this paper cites.
Image retrieval on real-life images with pre-trained vision-and-language models
Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould. 2021 · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021 · 2021
Earlier work this paper cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021 · 2021
Earlier work this paper cites.
Retrieval augmented code generation and summarization
Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Romain Vencu, Richard Beaumont, Robert Kaczmarczyk, Jenia Jitsev, Atsushi Komatsuzaki, et al. 2021 · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021 · 2021
Earlier work this paper cites.
Multimodal{qa}: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant. 2021 · 2021
Earlier work this paper cites.
The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. 2021 · 2021
Cited alongside, same era.
Webqa: Multihop and multimodal qa
Yingshan Chang, Mridu Narang, Hisami Suzuki, Guihong Cao, Jianfeng Gao, and Yonatan Bisk. 2022 · 2022
Cited alongside, same era.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William Cohen. 2022a · 2022
Cited alongside, same era.
Tpu-knn: K nearest neighbor search at peak flop/s
Felix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo, David Majnemer, and Sanjiv Kumar. 2022 · 2022
Cited alongside, same era.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2022 · 2022
Cited alongside, same era.
Robust multi model rag pipeline for documents containing text, table & images
Pankaj Joshi, Aditya Gupta, Pankaj Kumar, and Manas Sisodia. 2024 · 2024
Later among the works it cites.
An empirical comparison of video frame sampling methods for multi-modal rag retrieval
Mahesh Kandhare and Thibault Gisselbrecht. 2024 · 2024
Later among the works it cites.
Cadmr: Cross-attention and disentangled learning for multimodal recommender systems
Yasser Khalafaoui, Martino Lovisetto, Basarab Matei, and Nistor Grozavu. 2024 · 2024
Later among the works it cites.
RAGAR, your falsehood radar: RAG-augmented reasoning for political fact-checking using multimodal large language models
Mohammed Abdul Khaliq, Paul Yu-Chun Chang, Mingyang Ma, Bernhard Pflugfelder, and Filip Miletić. 2024 · 2024
Later among the works it cites.
Do you remember? dense video captioning with cross-modal memory retrieval
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi, and Seong Tae Kim. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022 · 2022
Cited alongside, same era.
Unsupervised dense information retrieval with contrastive learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022 · 2022
Cited alongside, same era.
ViQuAE, a dataset for knowledge-based visual question answering about named entities
Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, Jose G Moreno, and Jesús Lovón Melgarejo. 2022 · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022 · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Improving medical multi-modal contrastive learning with expert annotations
Yogesh Kumar and Pekka Marttinen. 2024 · 2024
Later among the works it cites.
Alzheimerrag: Multimodal retrieval augmented generation for pubmed articles
Aritra Kumar Lahiri and Qinmin Vivian Hu. 2024 · 2024
Later among the works it cites.
PlanRAG: A plan-then-retrieval augmented generation for generative large language models as decision makers
Myeonghwa Lee, Seonho An, and Min-Soo Kim. 2024 · 2024
Later among the works it cites.
RA-ISF: Learning to answer and understand from retrieval augmentation via iterative self-feedback
Yanming Liu, Xinyue Peng, Xuhong Zhang, Weihao Liu, Jianwei Yin, Jiannan Cao, and Tianyu Du. 2024c · 2024
Later among the works it cites.
Ovis: Structural embedding alignment for multimodal large language model
Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye. 2024 · 2024
Later among the works it cites.
How does the textual information affect the retrieval of multimodal in-context learning?
Yang Luo, Zangwei Zheng, Zirui Zhu, and Yang You. 2024a · 2024
Later among the works it cites.
Visually guided generative text-layout pre-training for document intelligence
Zhiming Mao, Haoli Bai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024 · 2024
Later among the works it cites.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D Plumbley, Yuexian Zou, and Wenwu Wang. 2024 · 2024
Later among the works it cites.
OMG-QA: Building open-domain multi-modal generative question answering systems
Linyong Nan, Weining Fang, Aylin Rasteh, Pouya Lahabi, Weijin Zou, Yilun Zhao, and Arman Cohan. 2024b · 2024
Later among the works it cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2024 · 2024
Later among the works it cites.
Enwar: A rag-empowered multi-modal llm framework for wireless environment perception
Ahmad M Nazar, Abdulkadir Celik, Mohamed Y Selim, Asmaa Abdallah, Daji Qiao, and Ahmed M Eltawil. 2024 · 2024
Later among the works it cites.
Multimodal learned sparse retrieval with probabilistic expansion control
Thong Nguyen, Mariya Hendriksen, Andrew Yates, and Maarten de Rijke. 2024 · 2024
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, and et al. 2024 · 2024
Later among the works it cites.
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, Jin Shi, Fan Wu, Pei Chu, Minghao Liu, Zhenxiang Li, Chao Xu, Bo Zhang, Botian Shi, Zhongying Tu, and Conghui He. 2024 · 2024
Later among the works it cites.
Graph retrieval-augmented generation for large language models: A survey
Tyler Thomas Procko and Omar Ochoa. 2024 · 2024
Later among the works it cites.
InFoBench: Evaluating instruction following ability in large language models
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024 · 2024
Later among the works it cites.
Raven: Multitask retrieval augmented vision-language learning
Varun Nagaraj Rao, Siddharth Choudhary, Aditya Deshpande, Ravi Kumar Satzoda, and Srikar Appalaraju. 2024 · 2024
Later among the works it cites.
Context embeddings for efficient answer generation in rag
David Rau, Shuai Wang, Hervé Déjean, and Stéphane Clinchant. 2024 · 2024
Later among the works it cites.
Beyond text: Optimizing rag with multimodal inputs for industrial applications
Monica Riedler and Stefan Langer. 2024 · 2024
Later among the works it cites.
Unirag: Universal retrieval augmentation for multi-modal large language models
Sahel Sharifymoghaddam, Shivani Upadhyay, Wenhu Chen, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
Contrastive transformer cross-modal hashing for video-text retrieval
Xiaobo Shen, Qianxin Huang, Long Lan, and Yuhui Zheng. 2024 · 2024
Later among the works it cites.
XL-HeadTags: Leveraging multimodal retrieval augmentation for the multilingual generation of news headlines and tags
Faisal Tareque Shohan, Mir Tafseer Nayeem, Samsul Islam, Abu Ubaida Akash, and Shafiq Joty. 2024 · 2024
Later among the works it cites.
Soccerrag: Multimodal soccer information retrieval via natural queries
Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, and Pål Halvorsen. 2024 · 2024
Later among the works it cites.
Surf: Teaching large vision-language models to selectively utilize retrieved information
Jiashuo Sun, Jihai Zhang, Yucheng Zhou, Zhaochen Su, Xiaoye Qu, and Yu Cheng. 2024a · 2024
Later among the works it cites.
Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami, Ryan A Rossi, and Dinesh Manocha. 2024 · 2024
Later among the works it cites.
Retrieval meets reasoning: Even high-school textbook knowledge benefits multimodal reasoning
Cheng Tan, Jingxuan Wei, Linzhuang Sun, Zhangyang Gao, Siyuan Li, Bihui Yu, Ruifeng Guo, and Stan Z. Li. 2024 · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul R. Barham, Tom Hennigan, Benjamin Lee, Fabio Viola, Malcolm Reynolds, Yuanzhong Xu, Ryan Doherty, Eli Collins, Clemens Meyer, Eliza Rutherford, Erica Moreira, Kareem Ayoub, Megha Goel, Jack Krawczyk, Cosmo Du, Ed Chi, Heng-Tze Cheng, Eric Ni, Purvi Shah, Patrick Kane, Betty Chan, Manaal Faruqui, Aliaksei Severyn, Hanzhao Lin, YaGuang Li, Yong Cheng, Abe Ittycheriah, Mahdis Mahdieh, Mia Chen, Pei Sun, Dustin Tran, Sumit Bagri, Balaji Lakshminarayanan, and et al. 2024 · 2024
Later among the works it cites.
Faster maximum inner product search in high dimensions
Mo Tiwari, Ryan Kang, Jaeyong Lee, Donghyun Lee, Christopher J Piech, Sebastian Thrun, Ilan Shomorony, and Martin Jinye Zhang. 2024 · 2024
Later among the works it cites.
A comprehensive survey of hallucination mitigation techniques in large language models
SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024 · 2024
Later among the works it cites.
Must: An effective and scalable framework for multimodal search of target modality
Mengzhao Wang, Xiangyu Ke, Xiaoliang Xu, Lu Chen, Yunjun Gao, Pinpin Huang, and Runkai Zhu. 2024c · 2024
Later among the works it cites.
Simple but effective raw-data level multimodal fusion for composed image retrieval
Haokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei, Liqiang Nie, and Tat-Seng Chua. 2024 · 2024
Later among the works it cites.
Synthetic multimodal question generation
Ian Wu, Sravan Jayanthi, Vijay Viswanathan, Simon Rosenberg, Sina Khoshfetrat Pakazad, Tongshuang Wu, and Graham Neubig. 2024a · 2024
Later among the works it cites.
RULE: Reliable multimodal RAG for factuality in medical vision language models
Peng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu, Yun Li, Gang Li, Linjun Zhang, and Huaxiu Yao. 2024b · 2024
Later among the works it cites.
Corrective retrieval augmented generation
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024 · 2024
Later among the works it cites.
Echosight: Advancing visual-language models with wiki knowledge
Yibin Yan and Weidi Xie. 2024 · 2024
Later among the works it cites.
Visrag: Vision-based retrieval-augmented generation on multi-modality documents
Shi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu, Shuo Wang, Xu Han, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Later among the works it cites.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd. 2024 · 2024
Later among the works it cites.
Self-adaptive multimodal retrieval-augmented generation
Wenjia Zhai. 2024 · 2024
Later among the works it cites.
OmAgent: A multi-modal agent framework for complex video understanding with task divide-and-conquer
Lu Zhang, Tiancheng Zhao, Heting Ying, Yibo Ma, and Kyusong Lee. 2024e · 2024
Later among the works it cites.
Unifashion: A unified vision-language model for multimodal fashion retrieval and generation
Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, and Xiao-Ming Wu. 2024 · 2024
Later among the works it cites.
Unirag: Unification, retrieval, and generation for multimodal question answering with pre-trained language models
Qi Zhi Lim, Chin Poo Lee, Kian Ming Lim, and Ahmad Kamsani Samingan. 2024 · 2024
Later among the works it cites.
Predicting micro-video popularity via multi-modal retrieval augmentation
Ting Zhong, Jian Lang, Yifan Zhang, Zhangtao Cheng, Kunpeng Zhang, and Fan Zhou. 2024 · 2024
Later among the works it cites.
Advanced embedding techniques in multimodal retrieval augmented generation a comprehensive study on cross modal ai applications
Ren Zhou. 2024 · 2024
Later among the works it cites.
Multilingual machine translation with large language models: Empirical results and analysis
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024a · 2024
Later among the works it cites.
Ufinebench: Towards text-based person retrieval with ultra-fine granularity
Jialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang, Tianyu Guo, Nong Sang, Yunhe Wang, and Changxin Gao. 2024 · 2024
Later among the works it cites.
Matryoshka multimodal models
Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. 2025 · 2025
Closest in time.
Kyoyun Choi, Byungmu Yoon, Soobum Kim, and Jonggwon Park. 2025 · 2025
Closest in time.
Mm-poisonrag: Disrupting multimodal rag with local and global poisoning attacks
Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios, Saikrishna Sanniboina, Nanyun Peng, Kai-Wei Chang, Daniel Kang, and Heng Ji. 2025 · 2025
Closest in time.
Ru-ai: A large multimodal dataset for machine-generated content detection
Liting Huang, Zhihao Zhang, Yiran Zhang, Xiyue Zhou, and Shoujin Wang. 2025 · 2025
Closest in time.
Videorag: Retrieval-augmented generation over video corpus
Soyeong Jeong, Kangsan Kim, Jinheon Baek, and Sung Ju Hwang. 2025 · 2025
Closest in time.
Mmad: A comprehensive benchmark for multimodal large language models in industrial anomaly detection
Xi Jiang, Jian Li, Hanqiu Deng, Yong Liu, Bin-Bin Gao, Yifeng Zhou, Jialin Li, Chengjie Wang, and Feng Zheng. 2025 · 2025
Closest in time.
Vr-rag: Open-vocabulary species recognition with rag-assisted large multi-modal models
Faizan Farooq Khan, Jun Chen, Youssef Mohamed, Chun-Mei Feng, and Mohamed Elhoseiny. 2025 · 2025
Closest in time.
Llave: Large language and vision embedding models with hardness-weighted contrastive learning
Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, and Jinsong Su. 2025 · 2025
Closest in time.
Retrieval-augmented dynamic prompt tuning for incomplete multimodal learning
Jian Lang, Zhangtao Cheng, Ting Zhong, and Fan Zhou. 2025 · 2025
Closest in time.
Drcap: Decoding clap latents with retrieval-augmented generation for zero-shot audio captioning
Xiquan Li, Wenxi Chen, Ziyang Ma, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Qiuqiang Kong, and Xie Chen. 2025c · 2025
Closest in time.
Speech retrieval-augmented generation without automatic speech recognition
Do June Min, Karel Mundnich, Andy Lapastora, Erfan Soltanmohammadi, Srikanth Ronanki, and Kyu Han. 2025 · 2025
Closest in time.
Cross-modal retrieval of chest x-ray images and diagnostic reports based on report entity graph and dual attention: Cross-modal retrieval of chest x-ray images and diagnostic reports…
Weihua Ou, Yingjie Chen, Linqing Liang, Jianping Gou, Jiahao Xiong, Jiacheng Zhang, Lingge Lai, and Lei Zhang. 2025 · 2025
Closest in time.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. 2025 · 2025
Closest in time.
Reasoning llms for user-aware multimodal conversational agents
Hamed Rahimi, Jeanne Cattoni, Meriem Beghili, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, and Mohamed Chetouani. 2025 · 2025
Closest in time.
Videorag: Retrieval-augmented generation with extreme long-context videos
Xubin Ren, Lingrui Xu, Long Xia, Shuaiqiang Wang, Dawei Yin, and Chao Huang. 2025 · 2025
Closest in time.
Fashion-rag: Multimodal fashion image editing via retrieval-augmented generation
Fulvio Sanguigni, Davide Morelli, Marcella Cornia, and Rita Cucchiara. 2025 · 2025
Closest in time.
Collex – a multimodal agentic rag system enabling interactive exploration of scientific collections
Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel, Martin Semmann, and Chris Biemann. 2025 · 2025
Closest in time.
Agentic retrieval-augmented generation: A survey on agentic rag
Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2025 · 2025
Closest in time.
How to bridge the gap between modalities: Survey on multimodal large language model
Shezheng Song, Xiaopeng Li, Shasha Li, Shan Zhao, Jie Yu, Jun Ma, Xiaoguang Mao, Weimin Zhang, and Meng Wang. 2025 · 2025
Closest in time.
Chunyu Sun, Bingyu Liu, Zhichao Cui, Anbin Qi, Tian hao Zhang, Dinghao Zhou, and Lewei Lu. 2025 · 2025
Closest in time.
Contextual asr with retrieval augmented large language model
Cihan Xiao, Zejiang Hou, Daniel Garcia-Romero, and Kyu J Han. 2025 · 2025
Closest in time.
Omgm: Orchestrate multiple granularities and modalities for efficient multimodal retrieval
Wei Yang, Jingjing Fu, Rui Wang, Jinyu Wang, Lei Song, and Jiang Bian. 2025 · 2025
Closest in time.
Woongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek, and Sung Ju Hwang. 2025 · 2025
Closest in time.
A multimodal multi-agent framework for radiology report generation
Ziruo Yi, Ting Xiao, and Mark V. Albert. 2025 · 2025
Closest in time.
Unveiling the potential of multimodal retrieval augmented generation with planning
Xiaohan Yu, Zhihan Yang, and Chong Chen. 2025 · 2025
Closest in time.
Murar: A simple and effective multimodal retrieval and answer refinement framework for multimodal question answering
Zhengyuan Zhu, Daniel Lee, Hong Zhang, Sai Sree Harsha, Loic Feujio, Akash Maharaj, and Yunyao Li. 2025 · 2025
Closest in time.