Fetching the paper…
Reading the bibliography…
We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions.
J. R. Dormand and P. J. Prince, “A family of embedded runge-kutta formulae,” Journal of computational and applied mathematics , vol. 6, no. 1, pp. 19–26, 1980
1980
Earlier work this paper cites.
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer, “Crowdmos: An approach for crowdsourcing mean opinion score studies,” in 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2011, pp. 2416–2419
2011
Earlier work this paper cites.
W. B. De Haas, J. P. Magalhães, and F. Wiering, “Improving audio chord transcription by exploiting harmonic and metric knowledge.” in ISMIR , 2012, pp. 295–300
2012
Earlier work this paper cites.
R. M. Bittner, B. McFee, J. Salamon, P. Q. Li, and J. P. Bello, “Deep salience representations for f0 estimation in polyphonic music,” in International Society for Music Information Retrieval Conference , 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:4531539
2017
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Mor, L. Wolf, A. Polyak, and Y. Taigman, “A universal music translation network,” 2018
2018
Earlier work this paper cites.
R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. Duvenaud, “Neural ordinary differential equations,” 2019
2019
Earlier work this paper cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , 2020
2020
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Earlier work this paper cites.
A. Défossez, “Hybrid spectrogram and waveform source separation,” in Proceedings of the ISMIR 2021 Workshop on Music Source Separation , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Rouard, F. Massa, and A. Défossez, “Hybrid transformers for music source separation,” 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” 2022
2022
Cited alongside, same era.
O. Press, N. A. Smith, and M. Lewis, “Train short, test long: Attention with linear biases enables input length extrapolation,” 2022
2022
Cited alongside, same era.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” 2022
2022
Cited alongside, same era.
T. Sugimoto, “Loudness-level-chasing algorithm for multiformat live audio production,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1290–1304, 2022
2022
Cited alongside, same era.
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” 2023
2023
Later among the works it cites.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Later among the works it cites.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu, “Voicebox: Text-guided multilingual universal speech generation at scale,” 2023
2023
Later among the works it cites.
A. Vyas, B. Shi, M. Le, A. Tjandra, Y.-C. Wu, B. Guo, J. Zhang, X. Zhang, R. Adkins, W. Ngan, J. Wang, I. Cruz, B. Akula, A. Akinyemi, B. Ellis, R. Moritz, Y. Yungster, A. Rakotoarison, L. Tan, C. Summers, C. Wood, J. Lane, M. Williamson, and W.-N. Hsu, “Audiobox: Unified audio generation with natural language prompts,” 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Forsgren and H. Martiros, “Riffusion-stable diffusion for real-time music generation. 2022,” URL https://riffusion. com/about
2022
Cited alongside, same era.
Y.-K. Wu, C.-Y. Chiu, and Y.-H. Yang, “Jukedrummer: Conditional beat-aware audio-domain drum accompaniment generation via transformer vq-vae,” 2022
2022
Cited alongside, same era.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank, “Musiclm: Generating music from text,” 2023
2023
Cited alongside, same era.
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, J. Engel, Q. V. Le, W. Chan, Z. Chen, and W. Han, “Noise2music: Text-conditioned music generation with diffusion models,” 2023
2023
Cited alongside, same era.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,” in International Conference on Machine Learning . PMLR, 2023, pp. 13 916–13 932
2023
Later among the works it cites.
C. Donahue, A. Caillon, A. Roberts, E. Manilow, P. Esling, A. Agostinelli, M. Verzetti, I. Simon, O. Pietquin, N. Zeghidour, and J. Engel, “Singsong: Generating musical accompaniments from singing,” 2023
2023
Later among the works it cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3836–3847
2023
Later among the works it cites.
M. Levy, B. D. Giorgi, F. Weers, A. Katharopoulos, and T. Nickson, “Controllable music production with diffusion models and guidance gradients,” 2023
2023
Later among the works it cites.
H. F. Garcia, P. Seetharaman, R. Kumar, and B. Pardo, “Vampnet: Music generation via masked acoustic token modeling,” 2023
2023
Later among the works it cites.
2024
Closest in time.
Z. Novack, J. McAuley, T. Berg-Kirkpatrick, and N. J. Bryan, “Ditto: Diffusion inference-time t-optimization for music generation,” 2024
2024
Closest in time.
J. D. Parker, J. Spijkervet, K. Kosta, F. Yesiler, B. Kuznetsov, J.-C. Wang, M. Avent, J. Chen, and D. Le, “Stemgen: A music generation model that listens,” 2024
2024
Closest in time.
L. Lin, G. Xia, Y. Zhang, and J. Jiang, “Arrange, inpaint, and refine: Steerable long-term music audio generation and editing via content-based controls,” 2024
2024
Closest in time.