Fetching the paper…
Reading the bibliography…
Audio is an essential part of our life, but creating it often requires expertise and is time-consuming.
Switchboard: Telephone speech corpus for research and development
J. J. Godfrey, E. C. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
A. Rix, J. Beerends, M. Hollier, and A. Hekstra · 2001
Earlier work this paper cites.
Fisher English training speech parts 1 and 2 LDC200{4,5}S13
Cieri, Christopher, et al. · 2004
Earlier work this paper cites.
Fisher English training speech parts 1 and 2 transcripts LDC200{4,5}T19
Cieri, Christopher, et al. · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger · 2017
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu · 2017
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
torchdiffeq, 2018
R. T. Q. Chen · 2018
Earlier work this paper cites.
Neural ordinary differential equations
R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud · 2018
Earlier work this paper cites.
C. Donahue, J. McAuley, and M. Puckette · 2018
Earlier work this paper cites.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2019
Earlier work this paper cites.
Hierarchical generative modeling for controllable speech synthesis
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, et al · 2019
Earlier work this paper cites.
Libri-Light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazar’e, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. rahman Mohamed, and E. Dupoux · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
C. D. Kim, B. Kim, H. Lee, and G. Kim · 2019
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
J. Yamagishi, C. Veaux, and K. MacDonald · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Cited alongside, same era.
A. Clifton, A. Pappu, S. Reddy, Y. Yu, J. Karlgren, B. Carterette, and R. Jones · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
J. Kong, J. Kim, and J. Bae · 2020
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2022
Later among the works it cites.
Musiclm: Generating music from text
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, et al · 2023
Closest in time.
Soundstorm: Efficient parallel audio generation
Z. Borsos, M. Sharifi, D. Vincent, E. Kharitonov, N. Zeghidour, and M. Tagliasacchi · 2023
Closest in time.
pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe
H. Bredin · 2023
Closest in time.
Simple and controllable music generation
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Improving transfer of expressivity for end-to-end multispeaker text-to-speech synthesis
A. Kulkarni, V. Colotte, and D. Jouvet · 2021
Cited alongside, same era.
Fastpitch: Parallel text-to-speech with pitch prediction
A. Łańcucki · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith, and M. Lewis · 2021
Cited alongside, same era.
The stable signature: Rooting watermarks in latent diffusion models
P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon · 2023
Closest in time.
Text-to-audio generation using instruction-tuned llm and latent diffusion model
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria · 2023
Closest in time.
PromptTTS: Controllable text-to-speech with text descriptions
Z. Guo, Y. Leng, Y. Wu, S. Zhao, and X. Tan · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision, 2023
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio
A. Kumar, K. Tan, Z. Ni, P. Manocha, X. Zhang, E. Henderson, and B. Xu · 2023
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale, 2023
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu · 2023
Closest in time.
Voiceldm: Text-to-speech with environmental context, 2023
Y. Lee, I. Yeon, J. Nam, and J. S. Chung · 2023
Closest in time.
PromptTTS 2: Describing and generating voices with text prompt
Y. Leng, Z. Guo, K. Shen, X. Tan, Z. Ju, Y. Liu, Y. Liu, D. Yang, L. Zhang, K. Song, et al · 2023
Closest in time.
Jen-1: Text-guided universal music generation with omnidirectional diffusion models
P. Li, B. Chen, Y. Yao, Y. Wang, A. Wang, and A. Wang · 2023
Closest in time.
Flow matching for generative modeling
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2023
Closest in time.
Expresso: A benchmark and analysis of discrete expressive speech resynthesis
T. A. Nguyen, W.-N. Hsu, A. d’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, et al · 2023
Closest in time.
Powerset multi-class cross entropy loss for neural speaker diarization
A. Plaquet and H. Bredin · 2023
Closest in time.
Mo \ \backslash ˆ usai: Text-to-music generation with long-context latent diffusion
F. Schneider, Z. Jin, and B. Schölkopf · 2023
Closest in time.
Seamless: Multilingual expressive and streaming speech translation
Seamless Communication · 2023
Closest in time.
Bespoke solvers for generative flow models
N. Shaul, J. Perez, R. T. Chen, A. Thabet, A. Pumarola, and Y. Lipman · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Z. Zhang, L. Zhou, C. Wang, S. Chen, Y. Wu, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Closest in time.