2023

MusicLM: Generating Music From Text

Agostinelli, Andrea, Denk, Timo I., Borsos, Zalán et al.

Understand

We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff".

  • MusicLM casts the process of conditional music generation as a hierarchical sequence-to-sequence modeling task, and it generates music at 24 kHz that remains consistent over several minutes.
  • Our experiments show that MusicLM outperforms previous systems both in audio quality and adherence to the text description.
  • Moreover, we demonstrate that MusicLM can be conditioned on both text and a melody in that it can transform whistled and hummed melodies according to the style described in a text caption.

Reading the bibliography…