Fetching the paper…
Reading the bibliography…
Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols.
G. E. Hinton and D. van Camp, “Keeping the neural networks simple by minimizing the description length of the weights,” in Annual Conference Computational Learning Theory , 1993
1993
Earlier work this paper cites.
C. M. Bishop, “Mixture density networks,” 1994
1994
Earlier work this paper cites.
P. Vincent, “A connection between score matching and denoising autoencoders,” Neural Computation , 2011
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Kraft and U. Zölzer, “BeaqleJS: HTML5 and JavaScript based framework for the subjective evaluation of audio quality,” in Linux Audio Conference , 2014. [Online]. Available: https://github.com/HSU-ANT/beaqlejs
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in ICML , 2015
2015
Earlier work this paper cites.
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, “Autoencoding beyond pixels using a learned similarity metric,” in ICML , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” NeurIPS , 2017
2017
Earlier work this paper cites.
X. Wang, S. Takaki, and J. Yamagishi, “An autoregressive recurrent mixture density network for parametric speech synthesis,” ICASSP , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Jin, G. J. Mysore, S. Diverdi, J. Lu, and A. Finkelstein, “Voco: Text-based insertion and replacement in audio narration,” TOG , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2017
2017
Earlier work this paper cites.
——, “The perception-distortion tradeoff,” in CVPR , 2018
2018
Earlier work this paper cites.
D. P. Kingma and P. Dhariwal, “Glow: Generative flow with invertible 1x1 convolutions,” NeurIPS , 2018
2018
Earlier work this paper cites.
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg et al. , “Parallel Wavenet: Fast high-fidelity speech synthesis,” in ICML , 2018
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” NeurIPS , 2018
2018
Earlier work this paper cites.
Y. Blau and T. Michaeli, “Rethinking lossy compression: The rate-distortion-perception tradeoff,” in ICML , 2019
2019
Earlier work this paper cites.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in ICLR , 2019
2019
Earlier work this paper cites.
B. Poole, S. Ozair, A. Van Den Oord, A. Alemi, and G. Tucker, “On variational bounds of mutual information,” in ICML , 2019
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Ping, K. Peng, K. Zhao, and Z. Song, “WaveFlow: A compact flow-based model for raw audio,” in ICML , 2020
2020
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Peng, W. Ping, Z. Song, and K. Zhao, “Non-autoregressive neural text-to-speech,” in ICML , 2020
2020
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed et al. , “On generative spoken language modeling from raw audio,” TACL , 2021
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” TASLP , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in CVPR , 2021
2021
Earlier work this paper cites.
G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mohamed, and B. Lakshminarayanan, “Normalizing flows for probabilistic modeling and inference,” JMLR , 2021
2021
Earlier work this paper cites.
R. J. Weiss, R. Skerry-Ryan, E. Battenberg, S. Mariooryad, and D. P. Kingma, “Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis,” in ICASSP , 2021
2021
Earlier work this paper cites.
C. Du and K. Yu, “Phone-level prosody modelling with gmm-based mdn for diverse and controllable speech synthesis,” TASLP , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Tan, L. Deng, Y. T. Yeung, X. Jiang, X. Chen, and T. Lee, “EditSpeech: A text based speech editing system using partial inference and bidirectional fusion,” in ASRU , 2021
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” in ICLR , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2023
Later among the works it cites.
H. Lu, G. Yang, N. Fei, Y. Huo, Z. Lu, P. Luo, and M. Ding, “VDT: General-purpose video diffusion transformers via mask modeling,” in ICLR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” in NeurIPS , 2022
2022
Cited alongside, same era.
T. Wang, J. Yi, L. Deng, R. Fu, J. Tao, and Z. Wen, “Context-aware mask prediction network for end-to-end text-based speech editing,” in ICASSP , 2022
2022
Cited alongside, same era.
H. Bai, R. Zheng, J. Chen, M. Ma, X. Li, and L. Huang, “A 3 T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,” in ICML , 2022
2022
Cited alongside, same era.
Z. Borsos, M. Sharifi, and M. Tagliasacchi, “SpeechPainter: Text-conditioned speech inpainting,” in Interspeech , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Shen, Z. Ju, X. Tan, E. Liu, Y. Leng, L. He, T. Qin, J. Bian et al. , “NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,” in ICLR , 2023
2023
Later among the works it cites.
Z. Jiang, Q. Yang, J. Zuo, Z. Ye, R. Huang, Y. Ren, and Z. Zhao, “FluentSpeech: Stutter-oriented automatic speech editing with context-aware diffusion models,” in ACL , 2023
2023
Later among the works it cites.
Z. Liu, Y. Guo, and K. Yu, “DiffVoice: Text-to-speech with latent diffusion,” in ICASSP , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency models,” in ICML , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Ye, W. Xue, X. Tan, J. Chen, Q. Liu, and Y. Guo, “CoMoSpeech: One-step speech and singing voice synthesis via consistency model,” in ACM Multimedia , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, “WeSpeaker: A research and production oriented speaker embedding learning toolkit,” in ICASSP , 2023
2023
Later among the works it cites.
J. Lovelace, S. Ray, K. Kim, K. Q. Weinberger, and F. Wu, “Simple-TTS: End-to-end text-to-speech synthesis with latent diffusion,” 2023. [Online]. Available: https://openreview.net/forum?id=m4mwbPjOwb
2023
Later among the works it cites.
2024
Closest in time.
M. Hassid, T. Remez, T. A. Nguyen, I. Gat, A. Conneau, F. Kreuk, J. Copet, A. Defossez, G. Synnaeve, E. Dupoux et al. , “Textually pretrained speech language models,” NeurIPS , 2024
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” NeurIPS , 2024
2024
Closest in time.
Z. Du, S. Zhang, K. Hu, and S. Zheng, “FunCodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,” in ICASSP , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Luo, T. Hu, S. Zhang, J. Sun, Z. Li, and Z. Zhang, “Diff-Instruct: A universal approach for transferring knowledge from pre-trained diffusion models,” NeurIPS , 2024
2024
Closest in time.
J.-Y. Franceschi, M. Gartrell, L. Dos Santos, T. Issenhuth, E. de Bézenac, M. Chen, and A. Rakotomamonjy, “Unifying GANs and score-based diffusion as generative particle models,” NeurIPS , 2024
2024
Closest in time.
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu, “ProlificDreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,” NeurIPS , 2024
2024
Closest in time.
2024
Closest in time.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , 2024
2024
Closest in time.
2024
Closest in time.
C. Du, Y. Guo, F. Shen, Z. Liu, Z. Liang, X. Chen, S. Wang, H. Zhang, and K. Yu, “UniCATS: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding,” in AAAI , 2024
2024
Closest in time.
2024
Closest in time.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al. , “VoiceBox: Text-guided multilingual universal speech generation at scale,” NeurIPS , 2024
2024
Closest in time.
S. Dieleman, “The paradox of diffusion distillation,” 2024. [Online]. Available: https://sander.ai/2024/02/28/paradox.html
2024
Closest in time.
2024
Closest in time.
W. Guan, Q. Su, H. Zhou, S. Miao, X. Xie, L. Li, and Q. Hong, “Reflow-TTS: A rectified flow model for high-fidelity text-to-speech,” in ICASSP , 2024
2024
Closest in time.
2024
Closest in time.
Y. A. Li, C. Han, V. Raghavan, G. Mischler, and N. Mesgarani, “StyleTTS 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,” NeurIPS , 2024
2024
Closest in time.
G. Cámbara, P. L. Tobing, M. Babianski, R. Vipperla, D. W. R. Shmelkin, G. Coccia, O. Angelini, A. Joly, M. Lajszczak, and V. Pollet, “Mapache: Masked parallel transformer for advanced speech editing and synthesis,” in ICASSP , 2024
2024
Closest in time.