Fetching the paper…
Reading the bibliography…
Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech.
H. Sacks, E. A. Schegloff, and G. Jefferson, “A simplest systematics for the organization of turn taking for conversation,” in Studies in the organization of conversational interaction . Elsevier, 1978, pp. 7–55
1978
Earlier work this paper cites.
H. Bredin, “pyannote. audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,” in Proc. INTERSPEECH . ISCA, 2023, pp. 1983–1987
1987
Earlier work this paper cites.
W. Ward, “Understanding spontaneous speech,” in Speech and Natural Language: Proceedings of a Workshop Held at Philadelphia, Pennsylvania, February 21-23, 1989 , 1989
1989
Earlier work this paper cites.
E. A. SCHEGLOFF, “Overlapping talk and the organization of turn-taking for conversation,” Language in Society , vol. 29, no. 1, p. 1–63, 2000
2000
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: A resource for the next generations of speech-to-text.” in LREC , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
P. Taylor, Text-to-speech synthesis . Cambridge university press, 2009
2009
Earlier work this paper cites.
M. Heldner and J. Edlund, “Pauses, gaps and overlaps in conversations,” Journal of Phonetics , vol. 38, no. 4, pp. 555–568, 2010
2010
Earlier work this paper cites.
M. Heldner, J. Edlund, and J. B. Hirschberg, “Pitch similarity in the vicinity of backchannels,” 2010
2010
Earlier work this paper cites.
N. Dethlefs, H. Hastie, H. Cuayáhuitl, Y. Yu, V. Rieser, and O. Lemon, “Information density and overlap in spoken dialogue,” Computer speech & language , vol. 37, pp. 82–97, 2016
2016
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in Proc. INTERSPEECH , 2017, pp. 498–502
2017
Earlier work this paper cites.
T. Nagata and H. Mori, “Defining laughter context for laughter synthesis with spontaneous speech corpus,” IEEE Transactions on Affective Computing , vol. 11, no. 3, pp. 553–559, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. v. Neumann, K. Kinoshita, M. Delcroix, S. Araki, T. Nakatani, and R. Haeb-Umbach, “All-neural online source separation, counting, and diarization for meeting analysis,” in Proc. ICASSP , 2019, pp. 91–95
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
S. C. Levinson, “On the human" interaction engine",” in Roots of human sociality . Routledge, 2020, pp. 39–69
2020
Earlier work this paper cites.
N. Tits, K. E. Haddad, and T. Dutoit, “Laughter Synthesis: Combining Seq2seq Modeling with Transfer Learning,” in Proc. INTERSPEECH , 2020, pp. 3401–3405
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, “Flow-tts: A non-autoregressive network for text to speech based on flow,” in Proc. ICASSP . IEEE, 2020, pp. 7209–7213
2020
Earlier work this paper cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” Advances in Neural Information Processing Systems , vol. 33, pp. 8067–8077, 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
E. Battenberg, R. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, “Location-relative attention mechanisms for robust long-form speech synthesis,” in Proc. ICASSP . IEEE, 2020, pp. 6194–6198
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in International Conference on Machine Learning . PMLR, 2021, pp. 8599–8608
2021
Cited alongside, same era.
T. A. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed et al. , “Generative spoken dialogue language modeling,” Transactions of the Association for Computational Linguistics , vol. 11, pp. 250–266, 2023
2023
Later among the works it cites.
K. Lee, K. Park, and D. Kim, “Dailytalk: Spoken dialogue dataset for conversational text-to-speech,” in Proc. ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Gillick, W. Deng, K. Ryokai, and D. Bamman, “Robust laughter detection in noisy environments.” in Proc. INTERSPEECH , 2021, pp. 2481–2485
2021
Cited alongside, same era.
L. Zhang, Z. Chen, and Y. Qian, “Enroll-aware attentive statistics pooling for target speaker verification,” Proc. INTERSPEECH , pp. 311–315, 2022
2022
Cited alongside, same era.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in International Conference on Machine Learning . PMLR, 2022, pp. 2709–2720
2022
Cited alongside, same era.
2022
Cited alongside, same era.
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “Prodiff: Progressive fast diffusion model for high-quality text-to-speech,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 2595–2605
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Liu, Y. Guo, and K. Yu, “Diffvoice: Text-to-speech with latent diffusion,” in Proc. ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
A. Larsen, S. Sønderby, H. Larochelle, and O. Winther, “Autoencoding beyond pixels using a learned similarity metric. december 31, 2015,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “AudioLM: a language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
C. Li, Y. Qian, Z. Chen, N. Kanda, D. Wang, T. Yoshioka, Y. Qian, and M. Zeng, “Adapting Multi-Lingual ASR Models for Handling Multiple Talkers,” in Proc. INTERSPEECH , 2023, pp. 1314–1318
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Zhang, Y. Qian, L. Yu, H. Wang, H. Yang, S. Liu, L. Zhou, and Y. Qian, “DDTSE: Discriminative diffusion model for target speech extraction,” IEEE Spoken Language Technology Workshop , 2024
2024
Closest in time.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al. , “Voicebox: Text-guided multilingual universal speech generation at scale,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
S. Kim, K. Shih, J. F. Santos, E. Bakhturina, M. Desta, R. Valle, S. Yoon, B. Catanzaro et al. , “P-flow: A fast and data-efficient zero-shot tts through speech prompting,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
S. Mehta, R. Tu, J. Beskow, É. Székely, and G. E. Henter, “Matcha-tts: A fast tts architecture with conditional flow matching,” in Proc. ICASSP . IEEE, 2024, pp. 11 341–11 345
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Closest in time.