Fetching the paper…
Reading the bibliography…
Speech pre-training has primarily demonstrated efficacy on classification tasks, while its capability of generating novel speech, similar to how GPT-2 can generate coherent paragraphs, has barely been explored.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Variation across speech and writing
Douglas Biber. 1991 · 1991
Earlier work this paper cites.
Prosody in the comprehension of spoken language: A literature review
Anne Cutler, Delphine Dahan, and Wilma Van Donselaar. 1997 · 1997
Earlier work this paper cites.
Can prosody aid the automatic classification of dialog acts in conversational speech?
Elizabeth Shriberg, Andreas Stolcke, Daniel Jurafsky, Noah Coccaro, Marie Meteer, Rebecca Bates, Paul Taylor, Klaus Ries, Rachel Martin, and Carol Van Ess-Dykema. 1998 · 1998
Earlier work this paper cites.
Prosody-based automatic segmentation of speech into sentences and topics
Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, and Gökhan Tür. 2000 · 2000
Earlier work this paper cites.
Prosodic features which cue back-channel responses in english and japanese
Nigel Ward and Wataru Tsukahara. 2000 · 2000
Earlier work this paper cites.
Prosody models for conversational speech recognition
Mari Ostendorf, Izhak Shafran, and Rebecca Bates. 2003 · 2003
Earlier work this paper cites.
Prosody modeling for automatic speech recognition and understanding
Elizabeth Shriberg and Andreas Stolcke. 2004 · 2004
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2006
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Using prosodic features in language models for meetings
Songfang Huang and Steve Renals. 2007 · 2007
Earlier work this paper cites.
Data augmenting contrastive learning of speech representations in the time domain
Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, and Emmanuel Dupoux. 2021 · 2007
Earlier work this paper cites.
A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments
Tomohiro Nakatani, Shigeaki Amano, Toshio Irino, Kentaro Ishizuka, and Tadahisa Kondo. 2008 · 2008
Earlier work this paper cites.
Exploiting prosodic breaks in language modeling with random forests
Yi Su and Frederick Jelinek. 2008 · 2008
Earlier work this paper cites.
Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend
Wei Chu and Abeer Alwan. 2009 · 2009
Earlier work this paper cites.
Unsupervised learning of disentangled speech content and style representation
Andros Tjandra, Ruoming Pang, Yu Zhang, and Shigeki Karita. 2020 · 2010
Cited alongside, same era.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Cited alongside, same era.
CROWDMOS: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer. 2011 · 2011
Cited alongside, same era.
DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization
Shaoshi Ling and Yuzong Liu. 2020 · 2012
Cited alongside, same era.
Prosodic and temporal features for language modeling for dialog
Nigel G Ward, Alejandro Vega, and Timo Baumann. 2012 · 2012
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Later among the works it cites.
Predicting glottal closure insufficiency using fundamental frequency contour analysis
Jacob T Cohen, Alma Cohen, Limor Benyamini, Yossi Adi, and Joseph Keshet. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Cited alongside, same era.
Ethnologue: Languages of the World, Nineteenth edition
M. Paul Lewis, Gary F. Simon, and Charles D. Fennig. 2016 · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al. 2016 · 2016
Cited alongside, same era.
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Cited alongside, same era.
Parsing speech: a neural approach to integrating lexical and acoustic-prosodic information
Trang Tran, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, and Mari Ostendorf. 2017 · 2017
Cited alongside, same era.
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Later among the works it cites.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn et al. 2020 · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
A. T. Liu, S. Yang, P. Chi, P. Hsu, and H. Lee. 2020 · 2020
Later among the works it cites.
Unsupervised cross-domain singing voice conversion
Adam Polyak, Lior Wolf, Yossi Adi, and Yaniv Taigman. 2020 · 2020
Later among the works it cites.
Towards unsupervised learning of speech features in the wild
Morgane Rivière and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021 · 2021
Closest in time.
Fastpitch: Parallel text-to-speech with pitch prediction
Adrian Łańcucki. 2021 · 2021
Closest in time.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Closest in time.
Blizzard Challenge 2013
SynSIG · 2021
Closest in time.
SUPERB: speech processing universal performance benchmark
Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee. 2021 · 2021
Closest in time.