2024

World Model on Million-Length Video And Language With Blockwise RingAttention

Liu, Hao, Yan, Wilson, Zaharia, Matei et al.

Understand

Enabling long-context understanding remains a key challenge in scaling existing sequence models -- a crucial component in developing generally intelligent models that can process and operate over long temporal horizons that potentially consist of millions of tokens.

  • In this paper, we aim to address these challenges by providing a comprehensive exploration of the full development process for producing 1M context language models and video-language models, setting new benchmarks in language retrieval and new capabilities in long video understanding.
  • We detail our long context data curation process, progressive context extension from 4K to 1M tokens, and present an efficient open-source implementation for scalable training on long sequences.
  • Additionally, we open-source a family of 7B parameter models capable of processing long text documents and videos exceeding 1M tokens.

Reading the bibliography…