2022

VarietySound: Timbre-Controllable Video to Sound Generation via Unsupervised Information Disentanglement

Cui, Chenye, Ren, Yi, Liu, Jinglin et al.

Understand

Video to sound generation aims to generate realistic and natural sound given a video input.

  • However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the generated sound timbre, leading to the problem that people cannot obtain the desired timbre under these methods sometimes.
  • In this paper, we pose the task of generating sound with a specific timbre given a video input and a reference audio sample.
  • To solve this task, we disentangle each target sound audio into three components: temporal information, acoustic information, and background information.

Reading the bibliography…