2020

Jukebox: A Generative Model for Music

Dhariwal, Prafulla, Jun, Heewoo, Payne, Christine et al.

Understand

We introduce Jukebox, a model that generates music with singing in the raw audio domain.

  • We tackle the long context of raw audio using a multi-scale VQ-VAE to compress it to discrete codes, and modeling those using autoregressive Transformers.
  • We show that the combined model at scale can generate high-fidelity and diverse songs with coherence up to multiple minutes.
  • We can condition on artist and genre to steer the musical and vocal style, and on unaligned lyrics to make the singing more controllable.

Reading the bibliography…