Fetching the paper…
Reading the bibliography…
Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“The zero resource speech challenge 2017,”
Ewan Dunbar, Xuan Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera, and Emmanuel Dupoux, · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Vqvae unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019,”
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura, · 2019
Earlier work this paper cites.
“The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,”
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, and Emmanuel Dupoux, · 2020
Earlier work this paper cites.
“Transformer vq-vae for unsupervised unit discovery and speech synthesis: Zerospeech 2020 challenge,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2020
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer,”
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu, · 2020
Earlier work this paper cites.
“On Generative Spoken Language Modeling from Raw Audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2021
Earlier work this paper cites.
“Text-free prosody-aware generative spoken language modeling,”
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al., · 2021
Earlier work this paper cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2021
Earlier work this paper cites.
“Textless speech emotion conversion using decomposed and discrete representations,”
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi, · 2021
Earlier work this paper cites.
“Direct speech-to-speech translation with discrete units,”
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al., · 2021
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Earlier work this paper cites.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al., · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Earlier work this paper cites.
“Slam: A unified encoder for speech and language modeling via speech-text joint pre-training,”
Ankur Bapna, Yu-an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, and Yu Zhang, · 2021
Earlier work this paper cites.
“Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,”
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, et al., · 2021
Cited alongside, same era.
“textless-lib: a library for textless spoken language processing,”
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, et al., · 2022
Cited alongside, same era.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour, · 2022
Cited alongside, same era.
“Contentvec: An improved self-supervised speech representation by disentangling speakers,”
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang, · 2022
Cited alongside, same era.
“Expresso: A benchmark and analysis of discrete expressive speech resynthesis,”
Tu Anh Nguyen, Wei-Ning Hsu, Antony d’Avirro, Bowen Shi, Itai Gat, Maryam Fazel-Zarani, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, et al., · 2023
Later among the works it cites.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al., · 2023
Later among the works it cites.
“Exploration of efficient end-to-end asr using discretized input from self-supervised learning,”
Xuankai Chang, Brian Yan, Yuya Fujita, Takashi Maekaku, and Shinji Watanabe, · 2023
Later among the works it cites.
Xuankai Chang, Brian Yan, Kwanghee Choi, Jeeweon Jung, Yichen Lu, Soumi Maiti, Roshan Sharma, Jiatong Shi, Jinchuan Tian, Shinji Watanabe, et al., · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al., · 2022
Cited alongside, same era.
“Speaking style conversion with discrete self-supervised units,”
Gallil Maimon and Yossi Adi, · 2022
Cited alongside, same era.
“Textless speech-to-speech translation on real data,”
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, and Wei-Ning Hsu, · 2022
Cited alongside, same era.
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee, · 2022
Cited alongside, same era.
Wei-Ning Hsu, Tal Remez, Bowen Shi, Jacob Donley, and Yossi Adi, · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Opt: Open pre-trained transformer language models,”
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al., · 2022
Cited alongside, same era.
“Mu 2 slam: Multitask, multilingual speech and language models,”
Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey, and Ankur Bapna, · 2022
Cited alongside, same era.
Later among the works it cites.
“Analysing discrete self supervised speech representation for spoken language modeling,”
Amitay Sicherman and Yossi Adi, · 2023
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,”
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour, · 2023
Later among the works it cites.
“Maestro-u: Leveraging joint speech-text representation learning for zero supervised speech asr,”
Zhehuai Chen, Ankur Bapna, Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Pedro Moreno, and Nanxin Chen, · 2023
Later among the works it cites.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al., · 2023
Later among the works it cites.
“Speechtokenizer: Unified speech tokenizer for speech large language models,”
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu, · 2023
Later among the works it cites.
“Audiotoken: Adaptation of text-conditioned diffusion models for audio-to-image generation,”
Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi, and Idan Schwartz, · 2023
Later among the works it cites.
Yifan Peng, Ilia Kulikov, Yilin Yang, Sravya Popuri, Hui Lu, Changhan Wang, and Hongyu Gong, · 2024
Closest in time.
“Textually pretrained speech language models,”
Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Defossez, Gabriel Synnaeve, Emmanuel Dupoux, et al., · 2024
Closest in time.
“Spoken question answering and speech continuation using spectrogram-powered llm,”
Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, RJ Skerry-Ryan, and Michelle Tadmor Ramanovich, · 2024
Closest in time.
“Nast: Noise aware speech tokenization for speech language models,”
Shoval Messica and Yossi Adi, · 2024
Closest in time.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2024
Closest in time.
“Diverse and aligned audio-to-video generation via text-to-video model adaptation,”
Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf, Idan Schwartz, and Yossi Adi, · 2024
Closest in time.