Fetching the paper…
Reading the bibliography…
Text-to-music generation models are now capable of generating high-quality music audio in broad styles.
S. S. Stevens, J. Volkmann, and E. B. Newman, “A scale for the measurement of the psychological magnitude pitch,” JASA , 1937
1937
Earlier work this paper cites.
M. E. Davies, N. Degara, and M. D. Plumbley, “Evaluation methods for musical audio beat tracking algorithms,” Queen Mary University of London Tech. Rep. C4DM-TR-09-06 , 2009
2009
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in ICLR , 2013
2013
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel, “MIR_EVAL: A transparent implementation of common mir metrics.” in ISMIR , 2014
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML , 2015
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015
2015
Earlier work this paper cites.
M. Müller, Fundamentals of music processing: Audio, analysis, algorithms, applications . Springer, 2015
2015
Earlier work this paper cites.
F. Krebs, S. Böck, and G. Widmer, “An efficient state-space model for joint tempo and meter tracking.” in ISMIR , 2015
2015
Earlier work this paper cites.
S. Böck, F. Korzeniowski, J. Schlüter, F. Krebs, and G. Widmer, “Madmom: A new python audio and music signal processing library,” in ACM International Conference on Multimedia , 2016
2016
Earlier work this paper cites.
S. Böck, F. Krebs, and G. Widmer, “Joint beat and downbeat tracking with recurrent neural networks.” in ISMIR , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in ICASSP , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals et al. , “Neural discrete representation learning,” NeurIPS , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Dieleman, A. van den Oord, and K. Simonyan, “The challenge of realistic music generation: modelling raw audio at scale,” NeurIPS , 2018
2018
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart et al. , “Efficient neural audio synthesis,” in ICML , 2018
2018
Earlier work this paper cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen et al. , “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in ICLR , 2019
2019
Earlier work this paper cites.
N. Mor, L. Wolf, A. Polyak, and Y. Taigman, “A universal music translation network,” in ICLR , 2019
2019
Earlier work this paper cites.
S. Huang, Q. Li, C. Anil, X. Bao, S. Oore, and R. B. Grosse, “TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) pipeline for musical timbre transfer,” in ICLR , 2019
2019
Cited alongside, same era.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,” in ICLR , 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
J. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable digital signal processing,” in ICLR , 2020
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , 2020
2020
Cited alongside, same era.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in ICASSP , 2022
2022
Later among the works it cites.
S. Forsgren and H. Martiros, “Riffusion: Stable diffusion for real-time music generation,” 2022. [Online]. Available: https://riffusion.com/about
2022
Later among the works it cites.
Y.-K. Wu, C.-Y. Chiu, and Y.-H. Yang, “JukeDrummer: conditional beat-aware audio-domain drum accompaniment generation via transformer VQ-VA,” in ISMIR , 2022
2022
Later among the works it cites.
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, and a. o. Dieleman, “General-purpose, long-context autoregressive modeling with Perceiver AR,” in ICML , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2020
2020
Cited alongside, same era.
P. Virtanen, , and SciPy 1.0 Contributors, “SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,” Nature Methods , 2020
2020
Cited alongside, same era.
K. Chen, C.-i. Wang, T. Berg-Kirkpatrick, and S. Dubnov, “Music SketchNet: Controllable music generation via factorized representations of pitch and rhythm,” in ISMIR , 2020
2020
Cited alongside, same era.
H. H. Tan and D. Herremans, “Music FaderNets: Controllable music generation based on high-level features via low-level feature modelling,” in ISMIR , 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” in NeurIPS Workshop on Deep Gen. Models and Downstream Applications , 2021
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in ICLR , 2021
2021
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” in ICML , 2023
2023
Closest in time.
2023
Closest in time.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in ICCV , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
B. McFee and et al., “librosa/librosa: 0.10.1,” Aug. 2023. [Online]. Available: https://doi.org/10.5281/zenodo.8252662
2023
Closest in time.
“DiffWave,” https://github.com/lmnt-com/diffwave , 2023
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo, “VampNet: Music generation via masked acoustic token modeling,” in ISMIR , 2023
2023
Closest in time.
S.-L. Wu and Y.-H. Yang, “MuseMorphose: Full-song and fine-grained piano music style transfer with one Transformer VAE,” IEEE/ACM TASLP , 2023
2023
Closest in time.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul et al. , “AudioLM: a language modeling approach to audio generation,” IEEE/ACM TASLP , 2023
2023
Closest in time.
D. Kim, Y. Kim, W. Kang, and I.-C. Moon, “Refining generative process with discriminator guidance in score-based diffusion models,” in ICML , 2023
2023
Closest in time.