Fetching the paper…
Reading the bibliography…
This work profoundly analyzes discrete self-supervised speech representations (units) through the eyes of Generative Spoken Language Modeling (GSLM).
“Voronoi diagrams—a survey of a fundamental geometric data structure,”
Franz Aurenhammer, · 1991
Earlier work this paper cites.
“Timit acoustic phonetic continuous speech corpus,”
John S Garofolo, · 1993
Earlier work this paper cites.
“V-measure: A conditional entropy-based external cluster evaluation measure,”
Andrew Rosenberg and Julia Hirschberg, · 2007
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov et al., · 2015
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen et al., · 2018
Earlier work this paper cites.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger et al., · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Earlier work this paper cites.
“Unsupervised pretraining transfers well across languages,”
Morgane Riviere et al., · 2020
Cited alongside, same era.
“Self-supervised contrastive learning for unsupervised phoneme segmentation,”
Felix Kreuk, Joseph Keshet, and Yossi Adi, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn et al., · 2020
Cited alongside, same era.
“Superb: Speech processing universal performance benchmark,”
Shu-wen Yang et al., · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Self-supervised speaker diarization,”
Yehoshua Dissen, Felix Kreuk, and Joseph Keshet, · 2022
Later among the works it cites.
“Generative spoken dialogue language modeling,”
Tu Anh Nguyen et al., · 2022
Later among the works it cites.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour, · 2022
Later among the works it cites.
“Phonetic analysis of self-supervised representations of english speech,”
Dan Wells, Hao Tang, and Korin Richmond, · 2022
Later among the works it cites.
“Probing phoneme, language and speaker information in unsupervised speech representations,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al., · 2021
Cited alongside, same era.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak et al., · 2021
Cited alongside, same era.
Maureen de Seyssel, Marvin Lavechin, Yossi Adi, Emmanuel Dupoux, and Guillaume Wisniewski, · 2022
Later among the works it cites.
“textless-lib: a library for textless spoken language processing,”
Eugene Kharitonov et al., · 2022
Later among the works it cites.
“On the robustness of self-supervised representations for spoken language modeling,”
Itai Gat, Felix Kreuk, Ann Lee, Jade Copet, Gabriel Synnaeve, Emmanuel Dupoux, and Yossi Adi, · 2022
Later among the works it cites.