Fetching the paper…
Reading the bibliography…
Large-scale generative models such as GPT and DALL-E have revolutionized the research community.
Switchboard: Telephone speech corpus for research and development
J. J. Godfrey, E. C. Holliman, and J. McDaniel · 1992
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
R. Kubichek · 1993
Earlier work this paper cites.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al · 2011
Earlier work this paper cites.
CrowdMOS: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
A regression approach to speech enhancement based on deep neural networks
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee · 2014
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger · 2017
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu · 2017
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Expressive speech synthesis via modeling expressions with variational autoencoder
K. Akuzawa, Y. Iwasawa, and Y. Matsuo · 2018
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
A. Brock, J. Donahue, and K. Simonyan · 2018
Earlier work this paper cites.
torchdiffeq, 2018
R. T. Q. Chen · 2018
Earlier work this paper cites.
Neural ordinary differential equations
R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu, et al · 2018
Earlier work this paper cites.
StarGAN-VC: non-parallel many-to-many voice conversion using star generative adversarial networks
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
D. P. Kingma and P. Dhariwal · 2018
Earlier work this paper cites.
The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. H. Kinnunen, and Z. Ling · 2018
Earlier work this paper cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous · 2018
Earlier work this paper cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Y. Wang, D. Stanton, Y. Zhang, R. J. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2019
Earlier work this paper cites.
Hierarchical generative modeling for controllable speech synthesis
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, et al · 2019
Earlier work this paper cites.
Libri-Light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazar’e, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. rahman Mohamed, and E. Dupoux · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2019
Earlier work this paper cites.
Sdr–half-baked or well done?
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey · 2019
Earlier work this paper cites.
SpecAugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Sequence-to-sequence modelling of F0 for speech emotion conversion
C. Robinson, N. Obin, and A. Roebel · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Cited alongside, same era.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
J. Yamagishi, C. Veaux, and K. MacDonald · 2019
Cited alongside, same era.
Libritts: A corpus derived from librispeech for text-to-speech
Fastspeech 2: Fast and high-quality end-to-end text to speech
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu · 2021
Later among the works it cites.
fairseq s 2 : A scalable and integrable speech synthesis toolkit
C. Wang, W.-N. Hsu, Y. Adi, A. Polyak, A. Lee, P.-J. Chen, J. Gu, and J. M. Pino · 2021
Later among the works it cites.
Artificial fingerprinting for generative models: Rooting deepfake attribution in training data
N. Yu, V. Skripniuk, S. Abdelnabi, and M. Fritz · 2021
Later among the works it cites.
XLS-R: self-supervised cross-lingual speech representation learning at scale
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli · 2022
Later among the works it cites.
A3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing
H. Bai, R. Zheng, J. Chen, X. Li, M. Ma, and L. Huang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. J. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Real time speech enhancement in the waveform domain
A. Défossez, G. Synnaeve, and Y. Adi · 2020
Cited alongside, same era.
ECAPA-TDNN: Emphasized Channel Attention, propagation and aggregation in TDNN based speaker verification
B. Desplanques, J. Thienpondt, and K. Demuynck · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Tts-portuguese corpus: a corpus for speech synthesis in brazilian portuguese
E. Casanova, A. C. Junior, C. Shulby, F. S. d. Oliveira, J. P. Teixeira, M. A. Ponti, and S. Aluísio · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, et al · 2022
Later among the works it cites.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Later among the works it cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Later among the works it cites.
W.-N. Hsu, T. Remez, B. Shi, J. Donley, and Y. Adi · 2022
Later among the works it cites.
FastDiff: A fast conditional diffusion model for high-quality speech synthesis
R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao · 2022
Later among the works it cites.
Textless speech emotion conversion using decomposed and discrete representations
F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T.-A. Nguyen, M. Rivière, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi · 2022
Later among the works it cites.
Generative spoken dialogue language modeling
T. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. M. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed, and E. Dupoux · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models
C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi · 2022
Later among the works it cites.
Universal speech enhancement with score-based diffusion
J. Serrà, S. Pascual, J. Pons, R. O. Araz, and D. Scaini · 2022
Later among the works it cites.
NaturalSpeech: End-to-end text to speech synthesis with human-level quality
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, X. Wang, Y. Leng, Y. Yi, L. He, F. K. Soong, T. Qin, S. Zhao, and T.-Y. Liu · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2022
Later among the works it cites.
Scaling laws for generative mixed-modal language models
A. Aghajanyan, L. Yu, A. Conneau, W.-N. Hsu, K. Hambardzumyan, S. Zhang, S. Roller, N. Goyal, O. Levy, and L. Zettlemoyer · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision, 2023
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
Flow matching for generative modeling
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z.-H. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei · 2023
Closest in time.