2019

All-in-One Image-Grounded Conversational Agents

Ju, Da, Shuster, Kurt, Boureau, Y-Lan et al.

Understand

As single-task accuracy on individual language and image tasks has improved substantially in the last few years, the long-term goal of a generally skilled agent that can both see and talk becomes more feasible to explore.

  • In this work, we focus on leveraging individual language and image tasks, along with resources that incorporate both vision and language towards that objective.
  • We design an architecture that combines state-of-the-art Transformer and ResNeXt modules fed into a novel attentive multimodal module to produce a combined model trained on many tasks.
  • We provide a thorough analysis of the components of the model, and transfer performance when training on one, some, or all of the tasks.

Reading the bibliography…