Fetching the paper…
Reading the bibliography…
Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text.
Dynamic Time Warping
Müller, M. 2007 · 2007
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Wavelets for intonation modeling in HMM speech synthesis
Suni, A.; Aalto, D.; Raitio, T.; Alku, P.; and Vainio, M. 2013 · 2013
Earlier work this paper cites.
Differences of pitch profiles in Germanic and slavic languages
Bistra, A.; Grazyna, D.; Bernd, M.; Frank, Z.; Jeanin, J.; and Magdalena, O.-P. 2014 · 2014
Earlier work this paper cites.
WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications
McFee; Brian; Raffel, C.; Liang, D.; Ellis, D. P.; McVicar, M.; Battenberg, E.; and Nieto, O. 2015 · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Metrics for Polyphonic Sound Event Detection
Mesaros, A.; Heittola, T.; and Virtanen, T. 2016 · 2016
Earlier work this paper cites.
WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications
Morise, M.; Yokomori, F.; and Ozawa, K. 2016 · 2016
Earlier work this paper cites.
Audio Set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P. W.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
Image-to-Image Translation with Conditional Adversarial Networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
van den Oord, A.; Vinyals, O.; and Kavukcuoglu, K. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation
Choi, Y.; Choi, M.-J.; Kim, M. S.; Ha, J.-W.; Kim, S.; and Choo, J. 2018 · 2018
Earlier work this paper cites.
AudioCaps: Generating Captions for Audios in The Wild
Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
Measuring a speaker’s acoustic correlates of pitch-but which? A contrastive analysis for perceived speaker charisma
Niebuhr, O.; and Skarnitzl, R. 2019 · 2019
Cited alongside, same era.
Generative Pretraining From Pixels
Chen, M.; Radford, A.; Child, R.; Wu, J.; Jun, H.; Luan, D.; and Sutskever, I. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Cited alongside, same era.
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
Kong, Q.; Cao, Y.; Iqbal, T.; Wang, Y.; Wang, W.; and Plumbley, M. D. 2020 · 2020
Cited alongside, same era.
Language Models are Few-Shot Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2020 · 2020
AudioGen: Textually Guided Audio Generation
Kreuk, F.; Synnaeve, A., Gabrieland Polyak; Singer, U.; Défossez, A.; Copet, J.; Parikh, D.; Taigman, Y.; and Adi, Y. 2022 · 2022
Later among the works it cites.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; Mcgrew, B.; Sutskever, I.; and Chen, M. 2022 · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; Parikh, D.; Gupta, S.; and Taigman, Y. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Benefit of Temporally-Strong Labels in Audio Event Classification
Hershey, S.; Ellis, D. P. W.; Fonseca, E.; Jansen, A.; Liu, C.; Channing Moore, R.; and Plakal, M. 2021 · 2021
Cited alongside, same era.
High-Resolution Complex Scene Synthesis with Transformers
Jahn, M.; Rombach, R.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; et al. 2021 · 2021
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
Audio-to-Image Cross-Modal Generation
Zelaszczyk, M.; and Mandziuk, J. 2021 · 2021
Cited alongside, same era.
Yang, D.; Yu, J.; Wang, H.; Wang, W.; Weng, C.; Zou, Y.; and Yu, D. 2022 · 2022
Later among the works it cites.
Text-to-Audio Generation using Instruction Tuned LLM and Latent Diffusion Model
Ghosal, D.; Majumder, N.; Mehrish, A.; and Poria, S. 2023 · 2023
Closest in time.
LatentKeypointGAN: Controlling GANs via Latent Keypoints
He, X.; Wandt, B.; and Rhodin, H. 2023 · 2023
Closest in time.
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
Huang, R.; Huang, J.; Yang, D.; Ren, Y.; Liu, L.; Li, M.; Ye, Z.; Liu, J.; Yin, X.; and Zhao, Z. 2023 · 2023
Closest in time.
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 · 2023
Closest in time.
Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M. D.; Zou, Y.; and Wang, W. 2023 · 2023
Closest in time.
Full-Band General Audio Synthesis with Score-Based Diffusion
Pascual, S.; Bhattacharya, G.; Yeh, C.; Pons, J.; and Serrà, J. 2023 · 2023
Closest in time.
Accommodating Audio Modality in CLIP for Multimodal Processing
Ruan, L.; Hu, A.; Song, Y.; Zhang, L.; Zheng, S.; and Jin, Q. 2023 · 2023
Closest in time.
PEER: A Collaborative Language Model
Schick, T.; Yu, J. A.; Jiang, Z.; Petroni, F.; Lewis, P.; Izacard, G.; You, Q.; Nalmpantis, C.; Grave, E.; and Riedel, S. 2023 · 2023
Closest in time.
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Wang, P.; Wang, S.; Lin, J.; Bai, S.; Zhou, X.; Zhou, J.; Wang, X.; and Zhou, C. 2023 · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.