Audiolm: a language modeling approach to audio generation
Original
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Later among the works it cites.
Audio retrieval with wavtext5k and clap training
Original
Deshmukh, S., Elizalde, B., and Wang, H · 2022
Later among the works it cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Original
Ding, M., Zheng, W., Hong, W., and Tang, J · 2022
Later among the works it cites.
Clap: Learning audio concepts from natural language supervision
Original
Elizalde, B., Deshmukh, S., Ismail, M. A., and Wang, H · 2022
Later among the works it cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Original
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Later among the works it cites.
Ssast: Self-supervised audio spectrogram transformer
Gong, Y., Lai, C.-I., Chung, Y.-A., and Glass, J · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Original
Ho, J. and Salimans, T · 2022
Later among the works it cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Original
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Later among the works it cites.
Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech synthesis
Original
Huang, R., Ren, Y., Liu, J., Cui, C., and Zhao, Z · 2022
Later among the works it cites.
Audio retrieval with natural language queries: A benchmark study
Koepke, A. S., Oncescu, A.-M., Henriques, J., Akata, Z., and Albanie, S · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Original
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Original
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Original
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Original
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Later among the works it cites.
Resolution-robust large mask inpainting with fourier convolutions
Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K., and Lempitsky, V · 2022
Later among the works it cites.
Masked autoencoders that listen
Original
Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., Feichtenhofer, C., et al · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Original
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.