Fetching the paper…
Reading the bibliography…
We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of metrics to automatically evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019 · 1911
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Data augmenting contrastive learning of speech representations in the time domain
Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, and Emmanuel Dupoux. 2021 · 2007
Earlier work this paper cites.
Natural language processing with python
Steven Bird, Edward Loper, and Ewan Klein. 2009 · 2009
Earlier work this paper cites.
Non-autoregressive predictive coding for learning speech representations from local dependencies
Alexander H Liu, Yu-An Chung, and James Glass. 2020 · 2011
Earlier work this paper cites.
CROWDMOS: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer. 2011 · 2011
Earlier work this paper cites.
Unsupervised speech recognition
Alexei Baevski, Wei-Ning Hsu, and Alexis Conneau. 2021 · 2012
Earlier work this paper cites.
Text-free image-to-speech synthesis using learned segmental units
Wei-Ning Hsu, David Harwath, Christopher Song, and James Glass. 2020 · 2012
Earlier work this paper cites.
A nonparametric Bayesian approach to acoustic model discovery
Chia-ying Lee and James Glass. 2012 · 2012
Earlier work this paper cites.
DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization
Shaoshi Ling and Yuzong Liu. 2020 · 2012
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
The Zero Resource Speech Challenge 2015: Proposed approaches and results
Maarten Versteegh, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux. 2016 · 2015
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly. 2016 · 2016
Earlier work this paper cites.
Variational inference for acoustic unit discovery
Lucas Ondel, Lukáš Burget, and Jan Černockỳ. 2016 · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Hidden Markov Model variational autoencoder for acoustic unit discovery
Janek Ebbers, Jahn Heymann, Lukas Drude, Thomas Glarner, Reinhold Haeb-Umbach, and Bhiksha Raj. 2017 · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. 2018 · 2018
Earlier work this paper cites.
Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner
Emmanuel Dupoux. 2018 · 2018
Earlier work this paper cites.
Full Bayesian Hidden Markov Model variational autoencoder for acoustic unit discovery
Thomas Glarner, Patrick Hanebrink, Janek Ebbers, and Reinhold Haeb-Umbach. 2018 · 2018
Cited alongside, same era.
FFTNet: A real-time speaker-dependent neural vocoder
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu. 2018 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Natural TTS synthesis by conditioning WaveNet on MEL spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu. 2018 · 2018
Cited alongside, same era.
Unsupervised acoustic unit representation learning for voice conversion using WaveNet auto-encoders
Mingjie Chen and Thomas Hain. 2020 · 2020
Later among the works it cites.
Improved speech representations with multi-target autoregressive predictive coding
Yu-An Chung and James Glass. 2020 · 2020
Later among the works it cites.
The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units
Ewan Dunbar, Julien Karadayi, Mathieu Bernard, Xuan-Nga Cao, Robin Algayres, Lucas Ondel, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
Libri-light: A benchmark for ASR with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux. 2020 · 2020
Later among the works it cites.
A convolutional deep Markov Model for unsupervised speech representation learning
Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, and James Glass. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Cited alongside, same era.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
The Zero Resource Speech Challenge 2019: TTS without T
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W. Black, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux. 2019 · 2019
Cited alongside, same era.
Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks
Ryan Eloff, André Nortje, Benjamin van Niekerk, Avashna Govender, Leanne Nortje, Arnu Pretorius, Elan van Biljon, Ewald van der Westhuizen, Lisa van Staden, and Herman Kamper. 2019 · 2019
Cited alongside, same era.
Combining adversarial training and disentangled speech representation for robust zero-resource subword modeling
Siyuan Feng, Tan Lee, and Zhiyuan Peng. 2019 · 2019
Cited alongside, same era.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Deep contextualized acoustic representations for semi-supervised speech recognition
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff. 2020 · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
A. T. Liu, S. Yang, P. Chi, P. Hsu, and H. Lee. 2020 · 2020
Later among the works it cites.
Exploring TTS without T using biologically/psychologically motivated neural network modules (ZeroSpeech 2020)
Takashi Morita and Hiroki Koda. 2020 · 2020
Later among the works it cites.
Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 Challenge
Benjamin van Niekerk, Leanne Nortje, and Herman Kamper. 2020 · 2020
Later among the works it cites.
Multi-task self-supervised learning for robust speech recognition
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio. 2020 · 2020
Later among the works it cites.
Towards unsupervised learning of speech features in the wild
Morgane Rivière and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
Transformer VQ-VAE for unsupervised unit discovery and speech synthesis: ZeroSpeech 2020 Challenge
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2020 · 2020
Later among the works it cites.
Cyclic spectral modeling for unsupervised unit discovery into voice conversion with excitation and waveform modeling
Patrick Lumban Tobing, Tomoki Hayashi, Yi-Chiao Wu, Kazuhiro Kobayashi, and Tomoki Toda. 2020 · 2020
Later among the works it cites.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
W. Wang, Q. Tang, and K. Livescu. 2020 · 2020
Later among the works it cites.
Self-supervised representations improve end-to-end speech translation
Anne Wu, Changhan Wang, Juan Pino, and Jiatao Gu. 2020 · 2020
Later among the works it cites.
Audio ALBERT: A lite BERT for self-supervised learning of audio representation
Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh, Shang-Wen Li, and Hung-yi Lee. 2021 · 2021
Closest in time.
HuBERT: How much can a bad teacher benefit ASR pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Closest in time.
Semi-supervised spoken language understanding via self-supervised speech and language model pretraining
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass. 2021 · 2021
Closest in time.
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, and Emmanuel Dupoux. 2020 · 2021
Closest in time.
Early phonetic learning without phonetic categories: Insights from large-scale simulations on realistic input
Thomas Schatz, Naomi H Feldman, Sharon Goldwater, Xuan-Nga Cao, and Emmanuel Dupoux. 2021 · 2021
Closest in time.