Fetching the paper…
Reading the bibliography…
This paper presents fairseq S^2, a fairseq extension for speech synthesis.
Semi-supervised sequence-to-sequence asr using unpaired speech and text
Murali Karthick Baskar, Shinji Watanabe, Ramon Astudillo, Takaaki Hori, Lukáš Burget, and Jan Černockỳ. 2019 · 1905
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Learning hierarchical discrete linguistic units from visually-grounded speech
David Harwath, Wei-Ning Hsu, and James Glass. 2019 · 1911
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim. 1984 · 1984
Earlier work this paper cites.
Using dynamic time warping to find patterns in time series
Donald J Berndt and James Clifford. 1994 · 1994
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Andrew J Hunt and Alan W Black. 1996 · 1996
Earlier work this paper cites.
Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds
Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne. 1999 · 1999
Earlier work this paper cites.
The htk book
Steve Young, Gunnar Evermann, Mark Gales, Thomas Hain, Dan Kershaw, Xunying Liu, Gareth Moore, Julian Odell, Dave Ollason, Dan Povey, et al. 2002 · 2002
Earlier work this paper cites.
Discretalk: Text-to-speech as a machine translation problem
Tomoki Hayashi and Shinji Watanabe. 2020 · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020a · 2006
Earlier work this paper cites.
Multispeech: Multi-speaker text to speech with transformer
Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, Tao Qin, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al. 2007 · 2007
Earlier work this paper cites.
The hmm-based speech synthesis system (hts) version 2.0
Heiga Zen, Takashi Nose, Junichi Yamagishi, Shinji Sako, Takashi Masuko, Alan W Black, and Keiichi Tokuda. 2007 · 2007
Earlier work this paper cites.
A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments
Tomohiro Nakatani et al. 2008 · 2008
Earlier work this paper cites.
Unsupervised cross-domain singing voice conversion
Adam Polyak, Lior Wolf, Yossi Adi, and Yaniv Taigman. 2020 · 2008
Earlier work this paper cites.
Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend
Wei Chu and Abeer Alwan. 2009 · 2009
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black. 2009 · 2009
Earlier work this paper cites.
Multilingual speech translation with efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2020 · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Flávio Ribeiro, Dinei Florêncio, Cha Zhang, and Michael Seltzer. 2011 · 2011
Earlier work this paper cites.
Text-free image-to-speech synthesis using learned segmental units
Wei-Ning Hsu, David Harwath, Christopher Song, and James Glass. 2020 · 2012
Earlier work this paper cites.
Phonemizer
Mathieu Bernard. 2015 · 2015
Earlier work this paper cites.
End-to-end text-dependent speaker verification
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer. 2016 · 2016
Cited alongside, same era.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa. 2016 · 2016
Cited alongside, same era.
Python interface to the webrtc voice activity detector
John Wiseman. 2016 · 2016
Cited alongside, same era.
Merlin: An open source neural network speech synthesis system
Zhizheng Wu, Oliver Watts, and Simon King. 2016 · 2016
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
Sercan Arik, Gregory Diamos, Andrew Gibiansky, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017 · 2017
Cited alongside, same era.
The lj speech dataset
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Later among the works it cites.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Later among the works it cites.
Facebook FAIR’s WMT19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Keith Ito and Linda Johnson. 2017 · 2017
Cited alongside, same era.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017 · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2017 · 2017
Cited alongside, same era.
Listening while speaking: Speech chain by deep learning
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Cited alongside, same era.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. 2017 · 2017
Cited alongside, same era.
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Later among the works it cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Speech recognition with augmented synthesized speech
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu. 2019 · 2019
Later among the works it cites.
VizSeq: a visual analysis toolkit for text generation tasks
Changhan Wang, Anirudh Jain, Danlu Chen, and Jiatao Gu. 2019 · 2019
Later among the works it cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020 · 2020
Later among the works it cites.
Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. 2020 · 2020
Later among the works it cites.
Real time speech enhancement in the waveform domain
Alexandre Defossez, Gabriel Synnaeve, and Yossi Adi. 2020 · 2020
Later among the works it cites.
Espnet-tts: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan. 2020 · 2020
Later among the works it cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Later among the works it cites.
Speech-to-speech translation without text
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2020 · 2020
Later among the works it cites.
Fairseq s2t: Fast speech-to-text modeling with fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. 2020 · 2020
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Closest in time.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021 · 2021
Closest in time.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Closest in time.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu. 2021 · 2021
Closest in time.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Ron J Weiss, RJ Skerry-Ryan, Eric Battenberg, Soroosh Mariooryad, and Diederik P Kingma. 2021 · 2021
Closest in time.
Denoispeech: Denoising text to speech with frame-level noise modeling
Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2021 · 2021
Closest in time.