2019

Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech

Harwath, David, Hsu, Wei-Ning, Glass, James

Understand

In this paper, we present a method for learning discrete linguistic units by incorporating vector quantization layers into neural models of visually grounded speech.

  • We show that our method is capable of capturing both word-level and sub-word units, depending on how it is configured.
  • What differentiates this paper from prior work on speech unit learning is the choice of training objective.
  • Rather than using a reconstruction-based loss, we use a discriminative, multimodal grounding objective which forces the learned units to be useful for semantic image retrieval.

Reading the bibliography…