2024

MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Ye, Hanrong, Zhang, Haotian, Daxberger, Erik et al.

Understand

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding.

  • To achieve this goal, we work on three fronts.
  • First, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for egocentric videos ranging from 30 seconds to one hour long in Ego4D based on human-annotated data.
  • This is one of the largest egocentric QA datasets.

Reading the bibliography…