Fetching the paper…
Reading the bibliography…
We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach.
“The origins of intelligence in children”
Jean Piaget and Margaret Cook · 1952
Earlier work this paper cites.
“In the blink of an eye: how vision sparked the big bang of evolution”, 2003
Andrew Parker · 2003
Earlier work this paper cites.
“TextCaps: a Dataset for Image Captioning with Reading Comprehension”, 2020
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach and Amanpreet Singh · 2003
Earlier work this paper cites.
“Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite”
Andreas Geiger, Philip Lenz and Raquel Urtasun · 2012
Earlier work this paper cites.
“Rich feature hierarchies for accurate object detection and semantic segmentation”
Ross Girshick, Jeff Donahue, Trevor Darrell and Jitendra Malik · 2014
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Compositional semantic parsing on semi-structured tables”
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
“Imagenet large scale visual recognition challenge”
Olga Russakovsky et al · 2015
Earlier work this paper cites.
“Sun rgb-d: A rgb-d scene understanding benchmark suite”
Shuran Song, Samuel Lichtenberg and Jianxiong Xiao · 2015
Earlier work this paper cites.
“A diagram is worth a dozen images”
Aniruddha Kembhavi et al · 2016
Earlier work this paper cites.
“Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations”
Ranjay Krishna et al · 2016
Earlier work this paper cites.
“Modeling Context in Referring Expressions”, 2016
Licheng Yu et al · 2016
Earlier work this paper cites.
“Visual7w: Grounded question answering in images”
Yuke Zhu, Oliver Groth, Michael Bernstein and Li Fei-Fei · 2016
Earlier work this paper cites.
“Making the v in vqa matter: Elevating the role of image understanding in visual question answering”
Yash Goyal et al · 2017
Earlier work this paper cites.
“Clevr: A diagnostic dataset for compositional language and elementary visual reasoning”
Justin Johnson et al · 2017
Earlier work this paper cites.
“Seq2sql: Generating structured queries from natural language using reinforcement learning”
Victor Zhong, Caiming Xiong and Richard Socher · 2017
Earlier work this paper cites.
“Don’t just assume; look and answer: Overcoming priors for visual question answering”
Aishwarya Agrawal, Dhruv Batra, Devi Parikh and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
“Vizwiz grand challenge: Answering visual questions from blind people”
Danna Gurari et al · 2018
Earlier work this paper cites.
“Dvqa: Understanding data visualizations via question answering”
Kushal Kafle, Brian Price, Scott Cohen and Christopher Kanan · 2018
Earlier work this paper cites.
“TallyQA: Answering complex counting questions”
Manoj Acharya, Kushal Kafle and Christopher Kanan · 2019
Earlier work this paper cites.
“Scene text visual question answering”
Ali Biten et al · 2019
Earlier work this paper cites.
“GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering”
Drew. Hudson and Christopher. Manning · 2019
Earlier work this paper cites.
“OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge”
Kenneth Marino, Mohammad Rastegari, Ali Farhadi and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
“OCR-VQA: Visual Question Answering by Reading Text in Images”, 2019
2019
Earlier work this paper cites.
“Towards vqa models that can read”
Amanpreet Singh et al · 2019
Earlier work this paper cites.
“Raven: A dataset for relational and analogical visual reasoning”
Chi Zhang et al · 2019
Earlier work this paper cites.
“Semantic understanding of scenes through the ade20k dataset”
Bolei Zhou et al · 2019
Earlier work this paper cites.
“nuscenes: A multimodal dataset for autonomous driving”
Holger Caesar et al · 2020
Earlier work this paper cites.
“Shortcut learning in deep neural networks”
Robert Geirhos et al · 2020
Earlier work this paper cites.
“PathVQA: 30000+ Questions for Medical Visual Question Answering”
Xuehai He et al · 2020
Earlier work this paper cites.
“The hateful memes challenge: Detecting hate speech in multimodal memes”
Douwe Kiela et al · 2020
Earlier work this paper cites.
“Connecting Vision and Language with Localized Narratives”
Jordi Pont-Tuset et al · 2020
Earlier work this paper cites.
“Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations”
Adel Ahmadyan et al · 2021
Earlier work this paper cites.
“ARKitScenes - A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D Data”
Gilad Baruch et al · 2021
Earlier work this paper cites.
“imagehash (fork)”
Johannes Buchner · 2021
Earlier work this paper cites.
“An empirical study of training self-supervised vision transformers”
Xinlei Chen, Saining Xie and Kaiming He · 2021
Earlier work this paper cites.
“Finqa: A dataset of numerical reasoning over financial data”
Zhiyu Chen et al · 2021
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2021
Earlier work this paper cites.
“AI2D-RST: A multimodal corpus of 1000 primary school science diagrams”
Tuomo Hiippala et al · 2021
Earlier work this paper cites.
“Perceiver: General perception with iterative attention”
Andrew Jaegle et al · 2021
Earlier work this paper cites.
“Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning”
Pan Lu et al · 2021
Earlier work this paper cites.
“Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning”
Pan Lu et al · 2021
Earlier work this paper cites.
“Docvqa: A dataset for vqa on document images”
Minesh Mathew, Dimosthenis Karatzas and CV Jawahar · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding”
Mike Roberts et al · 2021
Earlier work this paper cites.
“ALFWorld: Aligning Text and Embodied Environments for Interactive Learning”
Mohit Shridhar et al · 2021
Earlier work this paper cites.
“VisualMRC: Machine Reading Comprehension on Document Images”
Ryota Tanaka, Kyosuke Nishida and Sen Yoshida · 2021
Earlier work this paper cites.
“TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance”
Fengbin Zhu et al · 2021
Earlier work this paper cites.
“Flamingo: a visual language model for few-shot learning”
Jean-Baptiste Alayrac et al · 2022
Earlier work this paper cites.
“Latr: Layout-aware transformer for scene-text vqa”
Ali Biten et al · 2022
Earlier work this paper cites.
“HiTab: A hierarchical table dataset for question answering and natural language generation”
Zhoujun Cheng et al · 2022
Earlier work this paper cites.
“Masked autoencoders are scalable vision learners”
Kaiming He et al · 2022
Earlier work this paper cites.
“Screenqa: Large-scale question-answer pairs over mobile app screenshots”
Yu-Chung Hsiao, Fedir Zubach and Maria Wang · 2022
Earlier work this paper cites.
“Chart-to-text: A large-scale benchmark for chart summarization”
Shankar Kantharaj et al · 2022
Cited alongside, same era.
“Donut: Document understanding transformer without ocr”
Geewook Kim et al · 2022
Cited alongside, same era.
“A convnet for the 2020s”
Zhuang Liu et al · 2022
Cited alongside, same era.
“Learn to explain: Multimodal reasoning via thought chains for science question answering”
Pan Lu et al · 2022
Cited alongside, same era.
“Chartqa: A benchmark for question answering about charts with visual and logical reasoning”
Ahmed Masry et al · 2022
Cited alongside, same era.
“ChatGPT”, 2022
OpenAI · 2022
Cited alongside, same era.
“Pytorch fsdp: experiences on scaling fully sharded data parallel”
Yanli Zhao et al · 2023
Later among the works it cites.
“Don’t Make Your LLM an Evaluation Benchmark Cheater”
Kun Zhou et al · 2023
Later among the works it cites.
“Starling-7b: Improving llm helpfulness & harmlessness with rlaif”
Banghua Zhu et al · 2023
Later among the works it cites.
“Minigpt-4: Enhancing vision-language understanding with advanced large language models”
Deyao Zhu et al · 2023
Later among the works it cites.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long Ouyang et al · 2022
Cited alongside, same era.
“High-Resolution Image Synthesis With Latent Diffusion Models”
Robin Rombach et al · 2022
Cited alongside, same era.
“LLM Evals and Benchmarking”, 2022
Omar Sanseviero · 2022
Cited alongside, same era.
“Laion-5b: An open large-scale dataset for training next generation image-text models”
Christoph Schuhmann et al · 2022
Cited alongside, same era.
“A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge”
Dustin Schwenk et al · 2022
Cited alongside, same era.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei et al · 2022
Cited alongside, same era.
Hessa Alawwad et al · 2024
Closest in time.
“Probing the 3D Awareness of Visual Foundation Models”
Mohamed Banani et al · 2024
Closest in time.
“Automatikz: Text-guided synthesis of scientific vector graphics with tikz”
Jonas Belouadi, Anne Lauscher and Steffen Eger · 2024
Closest in time.
“Honeybee: Locality-enhanced projector for multimodal llm”
Junbum Cha, Wooyoung Kang, Jonghwan Mun and Byungseok Roh · 2024
Closest in time.
“Visually Dehallucinative Instruction Generation: Know What You Don’t Know”
Sungguk Cha, Jusung Lee, Younghyun Lee and Cheoljong Yang · 2024
Closest in time.
“A survey on evaluation of large language models”
Yupeng Chang et al · 2024
Closest in time.
“ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model”
Guiming Chen et al · 2024
Closest in time.
“Are We on the Right Way for Evaluating Large Vision-Language Models?”
Lin Chen et al · 2024
Closest in time.
“How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites”
Zhe Chen et al · 2024
Closest in time.
“Chatbot arena: An open platform for evaluating llms by human preference”
Wei-Lin Chiang et al · 2024
Closest in time.
“Mobilevlm v2: Faster and stronger baseline for vision language model”
Xiangxiang Chu et al · 2024
Closest in time.
“Instructblip: Towards general-purpose vision-language models with instruction tuning”
Wenliang Dai et al · 2024
Closest in time.
“Rlhf workflow: From reward modeling to online rlhf”
Hanze Dong et al · 2024
Closest in time.
“Data filtering networks”
Alex Fang et al · 2024
Closest in time.
“BLINK: Multimodal Large Language Models Can See but Not Perceive”
Xingyu Fu et al · 2024
Closest in time.
“Datacomp: In search of the next generation of multimodal datasets”, 2024
Samir Gadre et al · 2024
Closest in time.
“SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models”
Peng Gao et al · 2024
Closest in time.
“Prismatic vlms: Investigating the design space of visually-conditioned language models”
Siddharth Karamcheti et al · 2024
Closest in time.
“Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset”
Hugo Laurençon, Léo Tronchon and Victor Sanh · 2024
Closest in time.
“What matters when building vision-language models?”
Hugo Laurençon, Léo Tronchon, Matthieu Cord and Victor Sanh · 2024
Closest in time.
“LLaVA-NeXT: Stronger LLMs Supercharge Multimodal Capabilities in the Wild”, 2024
Bo Li et al · 2024
Closest in time.
“Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models”
Lei Li et al · 2024
Closest in time.
“Mini-gemini: Mining the potential of multi-modality vision language models”
Yanwei Li et al · 2024
Closest in time.
“LLaVA-NeXT: Improved reasoning, OCR, and world knowledge”, 2024
Haotian Liu et al · 2024
Closest in time.
“A Decade’s Battle on Dataset Bias: Are We There Yet?”
Zhuang Liu and Kaiming He · 2024
Closest in time.
“DeepSeek-VL: towards real-world vision-language understanding”
Haoyu Lu et al · 2024
Closest in time.
“Wizardcoder: Empowering code large language models with evol-instruct”
Ziyang Luo et al · 2024
Closest in time.
“OpenEQA: Embodied Question Answering in the Era of Foundation Models”
Arjun Majumdar et al · 2024
Closest in time.
“Mm1: Methods, analysis & insights from multimodal llm pre-training”
Brandon McKinzie et al · 2024
Closest in time.
“Orca-Math: Unlocking the potential of SLMs in Grade School Math”, 2024
Arindam Mitra, Hamed Khanpour, Corby Rosset and Ahmed Awadallah · 2024
Closest in time.
“gpt4o”, 2024
OpenAI · 2024
Closest in time.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2024
Closest in time.
“Design2Code: How Far Are We From Automating Front-End Engineering?”
Chenglei Si et al · 2024
Closest in time.
“Mass-producing failures of multimodal systems with language models”
Shengbang Tong, Erik Jones and Jacob Steinhardt · 2024
Closest in time.
“Eyes wide shut? exploring the visual shortcomings of multimodal llms”
Shengbang Tong et al · 2024
Closest in time.
“ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy”
Kirill Vishniakov, Zhiqiang Shen and Zhuang Liu · 2024
Closest in time.
“Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset”
Ke Wang et al · 2024
Closest in time.
“V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs”
Penghao Wu and Saining Xie · 2024
Closest in time.
“grok”, 2024
xAI · 2024
Closest in time.
“Demystifying clip data”
Hu Xu et al · 2024
Closest in time.
“Yi: Open foundation models by 01. ai”
Alex Young et al · 2024
Closest in time.
“Mammoth: Building math generalist models through hybrid instruction tuning”
Xiang Yue et al · 2024
Closest in time.
“Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi”
Xiang Yue et al · 2024
Closest in time.
“Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning”
Yuexiang Zhai et al · 2024
Closest in time.
“Investigating the catastrophic forgetting in multimodal large language models”
Yuexiang Zhai et al · 2024
Closest in time.
“Judging llm-as-a-judge with mt-bench and chatbot arena”
Lianmin Zheng et al · 2024
Closest in time.
“OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement”
Tianyu Zheng et al · 2024
Closest in time.