2018

Conditional End-to-End Audio Transforms

Haque, Albert, Guo, Michelle, Verma, Prateek

Understand

We present an end-to-end method for transforming audio from one style to another.

  • For the case of speech, by conditioning on speaker identities, we can train a single model to transform words spoken by multiple people into multiple target voices.
  • For the case of music, we can specify musical instruments and achieve the same result.
  • Architecturally, our method is a fully-differentiable sequence-to-sequence model based on convolutional and hierarchical recurrent neural networks.

Reading the bibliography…