Fetching the paper…
Reading the bibliography…
End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently.
Music, Imagination, and Culture
Nicholas Cook, · 1990
Earlier work this paper cites.
“Jukebox: A Generative Model for Music,”
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, · 2005
Earlier work this paper cites.
The construction and evaluation of statistical models of melodic structure in music perception and composition
M. T. Pearce, · 2005
Earlier work this paper cites.
“Expectation in melody: The influence of context and learning,”
M. Pearce and G. Wiggins, · 2006
Earlier work this paper cites.
“Complete MIDI 1.0 Detailed Specification,”
MIDI Manufacturers Association, · 2008
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio,”
A. vd. Oord · 2016
Earlier work this paper cites.
“Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity,”
E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux, · 2019
Earlier work this paper cites.
“Fréchet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,”
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, · 2019
Earlier work this paper cites.
“On the evaluation of generative models in music,”
L.-C. Yang and A. Lerch, · 2020
Earlier work this paper cites.
“SpecTNT: a Time-Frequency Transformer for Music Audio,”
W.-T. Lu and J.-C. Wang and M. Won and K. Choi and X. Song, · 2021
Cited alongside, same era.
“High Fidelity Neural Audio Compression,”
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, · 2022
Cited alongside, same era.
“Multi-instrument Music Synthesis with Spectrogram Diffusion,”
C. Hawthorne, I. Simon, A. Roberts, N. Zeghidour, J. Gardner, E. Manilow, and J. Engel, · 2022
Cited alongside, same era.
“CLAP: Learning Audio Concepts From Natural Language Supervision,”
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, · 2022
Cited alongside, same era.
“Classifier-Free Diffusion Guidance,”
J. Ho and T. Salimans, · 2022
Cited alongside, same era.
“Modeling Beats and Downbeats with a Time-Frequency Transformer,”
Y.-N. Hung, J.-C. Wang, X. Song, W.-T. Lu, and M. Won, · 2022
Cited alongside, same era.
“Simple and Controllable Music Generation,”
J. Copet · 2023
Closest in time.
“VampNet: Music Generation via Masked Acoustic Token Modeling,”
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo, · 2023
Closest in time.
“Noise2Music: Text-conditioned Music Generation with Diffusion Models,”
Q. Huang · 2023
Closest in time.
“Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion,”
F. Schneider, Z. Jin, and B. Schölkopf, · 2023
Closest in time.
“SingSong: Generating musical accompaniments from singing,”
C. Donahue · 2023
Closest in time.
“Multi-Source Diffusion Models for Simultaneous Music Generation and Separation,”
G. Mariani, I. Tallini, E. Postolache, M. Mancusi, L. Cosmo, and E. Rodolà, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“To Catch A Chorus, Verse, Intro, or Anything Else: Analyzing a Song with Structural Functions,”
J.-C. Wang, Y.-N. Hung, and J. B. L. Smith, · 2022
Cited alongside, same era.
“High-Fidelity Audio Compression with Improved RVQGAN,”
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, · 2023
Cited alongside, same era.
“MusicLM: Generating Music From Text,”
A. Agostinelli · 2023
Cited alongside, same era.
“SoundStorm: Efficient Parallel Audio Generation,”
Z. Borsos, M. Sharifi, D. Vincent, E. Kharitonov, N. Zeghidour, and M. Tagliasacchi, · 2023
Closest in time.
“LLaMA: Open and Efficient Foundation Language Models,”
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, · 2023
Closest in time.
“Multitrack Music Transcription with a Time-Frequency Perceiver,”
W.-T. Lu, J.-C. Wang, and Y.-N. Hung, · 2023
Closest in time.