2023

eP-ALM: Efficient Perceptual Augmentation of Language Models

Shukor, Mustafa, Dancette, Corentin, Cord, Matthieu

Understand

Large Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales.

  • On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best performance on challenging benchmarks.
  • With the abundance of such unimodal models, a natural question arises; do we need also to follow this trend to tackle multimodal tasks? In this work, we propose to rather direct effort to efficient adaptations of existing models, and propose to augment Language Models with perception.
  • Existing approaches for adapting pretrained models for vision-language tasks still rely on several key components that hinder their efficiency.

Reading the bibliography…