2023

A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Fu, Chaoyou, Zhang, Renrui, Wang, Zihan et al.

Understand

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry.

  • They endow Large Language Models (LLMs) with powerful capabilities in visual understanding, enabling them to tackle diverse multi-modal tasks.
  • Very recently, Google released Gemini, its newest and most capable MLLM built from the ground up for multi-modality.
  • In light of the superior reasoning capabilities, can Gemini challenge GPT-4V's leading position in multi-modal learning? In this paper, we present a preliminary exploration of Gemini Pro's visual understanding proficiency, which comprehensively covers four domains: fundamental perception, advanced cognition, challenging vision tasks, and various expert capacities.

Reading the bibliography…