Fetching the paper…
Reading the bibliography…
Multimodal Foundation Models (MMFMs) have demonstrated strong performance in both computer vision and natural language processing tasks.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
“FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents”
Guillaume Jaume, Hazim Kemal and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
“Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension”
Daesik Kim, Seonhoon Kim and Nojun Kwak · 2019
Earlier work this paper cites.
“TabFact: A Large-scale Dataset for Table-based Fact Verification”
Wenhu Chen et al · 2020
Earlier work this paper cites.
“IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning”
Pan Lu et al · 2021
Earlier work this paper cites.
“Spatial Dual-Modality Graph Reasoning for Key Information Extraction”, 2021
Hongbin Sun, Zhanghui Kuang, Xiaoyu Yue, Chenhao Lin and Wayne Zhang · 2021
Earlier work this paper cites.
“DocVQA: A Dataset for VQA on Document Images”
Minesh Mathew, Dimosthenis Karatzas and C.. Jawahar · 2021
Earlier work this paper cites.
“WebSRC: A Dataset for Web-Based Structural Reading Comprehension”
Xingyu Chen et al · 2021
Earlier work this paper cites.
“InfographicVQA”
Minesh Mathew et al · 2022
Earlier work this paper cites.
“Compositional Semantic Parsing with Large Language Models”, 2022
Andrew Drozdov et al · 2022
Cited alongside, same era.
“Transformer Language Models without Positional Encodings Still Learn Positional Information”
Adi Haviv, Ori Ram, Ofir Press, Peter Izsak and Omer Levy · 2022
Cited alongside, same era.
“Efficient Guided Generation for Large Language Models”, 2023
Brandon. Willard and Rémi Louf · 2023
Cited alongside, same era.
“In-Context Retrieval-Augmented Language Models”
Ori Ram et al · 2023
Cited alongside, same era.
“Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering”, 2023
Wenjin Wang, Yunhao Li, Yixin Ou and Yin Zhang · 2023
Cited alongside, same era.
“An Introduction to Vision-Language Modeling”, 2024
Florian Bordes et al · 2024
Closest in time.
Franz Cesista, Rui Aguiar, Jason Kim and Paolo Acilo · 2024
Closest in time.
“Hugginface Inference Endpoints Documentation” https://huggingface.co/docs/inference-endpoints/index [Accessed: 2024-06-09], 2024
Huggingface Team · 2024
Closest in time.
“Hugginface Text Generation Interface Documentation” https://huggingface.co/docs/text-generation-inference/index [Accessed: 2024-06-09], 2024
Huggingface Team · 2024
Closest in time.
“LLaVA-NeXT: Improved reasoning, OCR, and world knowledge”, 2024
Haotian Liu et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“Improved Baselines with Visual Instruction Tuning”, 2023
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Cited alongside, same era.
“DocILE Benchmark for Document Information Localization and Extraction”, 2023
Štěpán Šimsa et al · 2023
Cited alongside, same era.
“DocLLM: A layout-aware generative language model for multimodal document understanding”, 2023
Dongsheng Wang et al · 2023
Cited alongside, same era.
“Hermes 2 Pro - Mistral 7B” https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B [Accessed: 2024-06-09], 2024
interstellarninja, Teknium, theemozilla, karan4d and huemin · 2024
Closest in time.
“MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI”
Xiang Yue et al · 2024
Closest in time.
“Mistral 7B” https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.1 [Accessed: 2024-06-09], 2024
Mistral Team · 2024
Closest in time.
“Matryoshka Multimodal Models”, 2024
Mu Cai, Jianwei Yang, Jianfeng Gao and Yong Lee · 2024
Closest in time.