Fetching the paper…
Reading the bibliography…
Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder.
Multiple stage vector quantization for speech coding
Juang, B.-H. and Gray, A · 1982
Earlier work this paper cites.
Vector quantization
Gray, R. M · 1984
Earlier work this paper cites.
A review of vector quantization techniques
Vasuki, A. and Vanathi, P · 2006
Earlier work this paper cites.
Optimal transport: Old and new
Villani, C · 2009
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Ribeiro, F., Florêncio, D., Zhang, C., and Seltzer, M · 2011
Earlier work this paper cites.
A kernel two-sample test
Gretton, A., Bordwardt, K., Rasch, M., Schoelopf, B., and Smola, A · 2012
Earlier work this paper cites.
Speechtokenizer: Unified speech tokenizer for speech large language models
Zhang, X., Zhang, D., Li, S., Zhou, Y., and Qiu, X · 2012
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, F., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. and Welling, M · 2014
Earlier work this paper cites.
An alternative update rule for generative adversarial networks
Huszar, F · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Wassertein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
Learning independent features with adversarial nets for non-linear ICA
Brakel, P. and Bengio, Y · 2017
Earlier work this paper cites.
Understanding disentangling in β \beta -VAE
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A · 2017
Earlier work this paper cites.
β \beta -vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Algorithms to measure audio programme loudness and true-peak audio level
ITU-R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
MINE: Mutual Information Neural Estimation
Belghazi, M. I., Baratin, A., Rajeswar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, R. D · 2018
Cited alongside, same era.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Cited alongside, same era.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brebisson, A., Bengio, Y., and Courville, A · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
MusicLM: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C · 2023
Later among the works it cites.
AudioLM: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2023
Later among the works it cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Later among the works it cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Cited alongside, same era.
Huang, Q., Park, D. S., Wang, T., Denk, T. I., Ly, A., Chen, N., Zhang, Z., Zhang, Z., Yu, J., Frank, C., Engel, J., Le, Q. V., Chan, W., Chen, Z., and Han, W · 2023
Later among the works it cites.
Nonlinear Independent Component Analysis for Principled Disentanglement in Unsupervised Deep Learning
Hyvarinen, A., Khemakhem, I., and Morioka, H · 2023
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2023
Later among the works it cites.
Deep deterministic independent component analysis for hyperspectral unmixing
Li, H., YU, S., and Principe, J · 2023
Later among the works it cites.
Mustango: Toward controllable text-to-music generation
Melechovsky, J., Guo, Z., Ghosal, D., Majumder, N., Herremans, D., and Poria, S · 2023
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F · 2023
Later among the works it cites.
Adapting frechet audio distance for generative music evaluation
Gui, A., Gamper, H., Braun, S., and Emmanouilidou, D · 2024
Closest in time.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Ju, Z., Wang, Y., Shen, K., Tan, X., Xin, D., Yang, D., Liu, Y., Leng, Y., Song, K., Tang, S., Wu, Z., Qin, T., Li, X.-Y., Ye, W., Zhang, S., Bian, J., He, L., Li, J., and Zhao, S · 2024
Closest in time.
High-fidelity audio compression with improved rvqgan, 2024
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2024
Closest in time.