Fetching the paper…
Reading the bibliography…
We introduce ResGen, an efficient Residual Vector Quantization (RVQ)-based generative model for high-fidelity generation with fast sampling.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Kim, J., Kim, S., Kong, J., and Yoon, S · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R · 2021
Earlier work this paper cites.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Casanova, E., Weber, J., Shulby, C. D., Junior, A. C., Gölge, E., and Ponti, M. A · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
Vector quantized diffusion model for text-to-image synthesis
Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B · 2022
Earlier work this paper cites.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Earlier work this paper cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Earlier work this paper cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Cited alongside, same era.
Soundstorm: Efficient parallel audio generation
Borsos, Z., Sharifi, M., Vincent, D., Kharitonov, E., Zeghidour, N., and Tagliasacchi, M · 2023
Cited alongside, same era.
Patched denoising diffusion models for high-resolution image synthesis
Ding, Z., Zhang, M., Wu, J., and Tu, Z · 2023
Cited alongside, same era.
Scalecrafter: Tuning-free higher-resolution visual generation with diffusion models
He, Y., Yang, S., Chen, H., Cun, X., Xia, M., Zhang, Y., Wang, X., He, R., Chen, Q., and Shan, Y · 2023
Cited alongside, same era.
Lumiere: A space-time diffusion model for video generation
Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Liu, G., Raj, A., et al · 2024
Closest in time.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2024
Closest in time.
Fast timing-conditioned latent audio diffusion
Evans, Z., Carr, C., Taylor, J., Hawley, S. H., and Pons, J · 2024
Closest in time.
Clam-tts: Improving neural codec language model for zero-shot text-to-speech
Kim, J., Lee, K., Chung, S., and Cho, J · 2024
Closest in time.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Closest in time.
Finite scalar quantization: Vq-vae made simple
Mentzer, F., Minnen, D., Agustsson, E., and Tschannen, M · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling up gans for text-to-image synthesis
Kang, M., Zhu, J.-Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T · 2023
Cited alongside, same era.
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Cited alongside, same era.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., and Hsu, W.-N · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Audiopalm: A large language model that can speak and listen
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., Quitry, F. d. C., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Shen, K., Ju, Z., Tan, X., Liu, E., Leng, Y., He, L., Qin, T., Bian, J., et al · 2024
Closest in time.
Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Siuzdak, H · 2024
Closest in time.
Hart: Efficient visual generation with hybrid autoregressive transformer
Tang, H., Wu, Y., Yang, S., Xie, E., Chen, J., Chen, J., Zhang, Z., Cai, H., Lu, Y., and Han, S · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Closest in time.
Givt: Generative infinite-vocabulary transformers
Tschannen, M., Eastwood, C., and Mentzer, F · 2024
Closest in time.
DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors
Lee, K., Kim, D. W., Kim, J., Chung, S., and Cho, J · 2025
Closest in time.