2022

Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

Wu, Nan, Jastrzębski, Stanisław, Cho, Kyunghyun et al.

Understand

We hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities.

  • Such behavior is counter-intuitive and hurts the models' generalization, as we observe empirically.
  • To estimate the model's dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality.
  • We refer to this gain as the conditional utilization rate.

Reading the bibliography…