Fetching the paper…
Reading the bibliography…
Text-to-audio (TTA) system has recently gained attention for its ability to synthesize general audio based on text descriptions.
A new way in sound synthesis
Andresen, U · 1979
Earlier work this paper cites.
Digital synthesis of plucked-string and drum timbres
Karplus, K. and Strong, A · 1983
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Salamon, J., Jacoby, C., and Bello, J. P · 2014
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
AudioSet: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2017
Earlier work this paper cites.
Audio super resolution using neural networks
Kuleshov, V., Enam, S. Z., and Ermon, S · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
CSTR VCTK corpus: English multi-speaker corpus for cstr voice cloning toolkit
Yamagishi, J., Veaux, C., MacDonald, K., et al · 2019
Earlier work this paper cites.
Clotho: an audio captioning dataset
Drossos, K., Lipping, S., and Virtanen, T · 2020
Earlier work this paper cites.
Ddsp: Differentiable digital signal processing
Engel, J., Hantrakul, L., Gu, C., and Roberts, A · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Cited alongside, same era.
Wavegrad: Estimating gradients for waveform generation
Chen, N., Zhang, Y., Zen, H., Weiss, R., Norouzi, M., and Chan, W · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
Lsgm: Score-based generative modeling in latent space
Vahdat, A., Kreis, K., and Kautz, J · 2021
Later among the works it cites.
Towards robust speech super-resolution
Wang, H. and Wang, D · 2021
Later among the works it cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Later among the works it cites.
Bilateral denoising diffusion models
Lam, M., Wang, J., Huang, R., Su, D., and Yu, D · 2022
Later among the works it cites.
Priorgrad: Improving conditional denoising diffusion models with data-driven adaptive prior
Lee, S., Kim, H., Shin, C., Tan, X., Liu, C., Meng, Q., Qin, T., Chen, W., Yoon, S., and Liu, T · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PSLA: Improving audio tagging with pretraining, sampling, labeling, and aggregation
Gong, Y., Chung, Y.-A., and Glass, J · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. and Dhariwal, P · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Cited alongside, same era.
Grad-tts: A diffusion probabilistic model for text-to-speech
Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Later among the works it cites.
Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis
Leng, Y., Chen, Z., Guo, J., Liu, H., Chen, J., Tan, X., Mandic, D., He, L., Li, X.-Y., Qin, T., et al · 2022
Later among the works it cites.
Full-band general audio synthesis with score-based diffusion
Pascual, S., Bhattacharya, G., Yeh, C., Pons, J., and Serrà, J · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2022
Later among the works it cites.
NaturalSpeech: End-to-end text to speech synthesis with human-level quality
Tan, X., Chen, J., Liu, H., Cong, J., Zhang, C., Liu, Y., Wang, X., Leng, Y., Yi, Y., He, L., et al · 2022
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S · 2022
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D · 2022
Later among the works it cites.
Audio-to-image cross-modal generation
Żelaszczyk, M. and Mańdziuk, J · 2022
Later among the works it cites.